# Veo 3.1 Text to Video > Create cinematic 8-second videos with Veo 3.1, Google’s latest text-to-video model in the Gemini API — now with native audio, frame control, and reference image support. ## Overview - **Endpoint**: `https://queue.modelrunner.run/google/veo-3.1/text-to-video` - **Model ID**: `google/veo-3.1/text-to-video` - **Category**: text-to-video - **Kind**: inference - **Tags**: video, text to video ## Pricing - **Price**: $0.4 per output second ## Request Lifecycle This model runs on the ModelRunner **asynchronous queue API** — a single POST does not return the output. Every call requires an `Authorization: Key $MODEL_RUNNER_KEY` header. Run three steps: 1. **Submit** — `POST https://queue.modelrunner.run/google/veo-3.1/text-to-video` with a JSON body holding the input fields at the top level. The body may also include a reserved top-level `metadata` object — a flat string map (max 16 keys, key ≤64 / value ≤512 chars) stored on the request for your own tagging. It is never sent to the model; filter your request history with `GET https://queue.modelrunner.run/requests?metadata=` (exact key=value matches, AND-ed). The response carries request handles only (no output yet): ```json { "status": "IN_QUEUE", "request_id": "<21-char id>", "status_url": "https://queue.modelrunner.run/google/veo-3.1/text-to-video/requests//status", "response_url": "https://queue.modelrunner.run/google/veo-3.1/text-to-video/requests/", "cancel_url": "https://queue.modelrunner.run/google/veo-3.1/text-to-video/requests//cancel" } ``` 2. **Poll status** — `GET ` until `status` is `COMPLETED`. Possible values are `IN_QUEUE`, `IN_PROGRESS`, `COMPLETED`, `FAILED`, `CANCELLED`. A `FAILED` request responds with HTTP 400 and an `error` field. 3. **Read result** — `GET `. Returns the finished request, including the generated `output`: ```json { "id": "", "status": "COMPLETED", "output": ..., "input": ... } ``` The JavaScript and Python SDKs below perform steps 2–3 for you. In any language without an SDK (Swift, Go, Kotlin, etc.) you must implement the polling loop and the final result fetch yourself — see the cURL example for the full flow. ### Input Schema - **`prompt`** (`string`, _required_): Text description of the desired video. Supports cinematic and natural language prompts. - **`duration`** (`duration`, _optional_): Video duration in seconds. - Default: `8` - Options: `4`, `6`, `8` - **`resolution`** (`resolution`, _optional_): Video resolution. Use 1080p for higher fidelity. - Default: `"720p"` - Options: `"720p"`, `"1080p"` - **`aspect_ratio`** (`aspect_ratio`, _optional_): Aspect ratio of the generated video. - Default: `"16:9"` - Options: `"16:9"`, `"9:16"` - **`negative_prompt`** (`string`, _optional_): Optional text describing what should not appear in the video. - **`person_generation`** (`person_generation`, _optional_): Controls whether people may be generated. Only 'allow_all' supported for text-to-video. - Default: `"allow_all"` - Options: `"allow_all"` ### Output Schema _No `Output` schema properties are available._ ## Default Example **Input** ```json { "prompt": "An eye-level shot glides through a misty pine forest at dawn. Soft sunlight filters through the trees, illuminating particles in the air. A red fox slowly emerges from the fog, stretches, and walks across moss-covered ground. The camera tracks its gentle movement in shallow focus. Natural ambiance fills the soundscape — birds chirping, distant rustle of leaves, and a light breeze passing through the forest.", "duration": 6, "resolution": "720p", "aspect_ratio": "16:9", "person_generation": "allow_all" } ``` **Output** ```json [ "https://media.modelrunner.ai/eBnC2hvJGoVrAoNcAniDW.mp4" ] ``` ## Usage Examples ### cURL The queue API is asynchronous: submit the request, poll `status_url` until it is `COMPLETED`, then read the result from `response_url`. Requires `jq`. ```bash # 1. Submit the request (returns request handles, not the output) SUBMIT=$(curl --silent --request POST \ --url https://queue.modelrunner.run/google/veo-3.1/text-to-video \ --header "Authorization: Key $MODEL_RUNNER_KEY" \ --header "Content-Type: application/json" \ --data '{ "prompt": "An eye-level shot glides through a misty pine forest at dawn. Soft sunlight filters through the trees, illuminating particles in the air. A red fox slowly emerges from the fog, stretches, and walks across moss-covered ground. The camera tracks its gentle movement in shallow focus. Natural ambiance fills the soundscape — birds chirping, distant rustle of leaves, and a light breeze passing through the forest.", "duration": 6, "resolution": "720p", "aspect_ratio": "16:9", "person_generation": "allow_all" }') STATUS_URL=$(echo "$SUBMIT" | jq -r '.status_url') RESPONSE_URL=$(echo "$SUBMIT" | jq -r '.response_url') # 2. Poll until the request leaves the queue / in-progress state while true; do STATUS=$(curl --silent --url "$STATUS_URL" \ --header "Authorization: Key $MODEL_RUNNER_KEY" | jq -r '.status') echo "Status: $STATUS" case "$STATUS" in COMPLETED) break ;; FAILED|CANCELLED) echo "Request $STATUS"; exit 1 ;; esac sleep 1 done # 3. Read the finished request, including the generated output curl --silent --url "$RESPONSE_URL" \ --header "Authorization: Key $MODEL_RUNNER_KEY" ``` ### JavaScript ```javascript import { modelrunner } from "@modelrunner/client"; const result = await modelrunner.subscribe("google/veo-3.1/text-to-video", { input: { "prompt": "An eye-level shot glides through a misty pine forest at dawn. Soft sunlight filters through the trees, illuminating particles in the air. A red fox slowly emerges from the fog, stretches, and walks across moss-covered ground. The camera tracks its gentle movement in shallow focus. Natural ambiance fills the soundscape — birds chirping, distant rustle of leaves, and a light breeze passing through the forest.", "duration": 6, "resolution": "720p", "aspect_ratio": "16:9", "person_generation": "allow_all" } }); console.log(result.data); ``` ### Python ```python import asyncio import modelrunner_ai async def main(): response = await modelrunner_ai.submit_async( "google/veo-3.1/text-to-video", arguments={ "prompt": "An eye-level shot glides through a misty pine forest at dawn. Soft sunlight filters through the trees, illuminating particles in the air. A red fox slowly emerges from the fog, stretches, and walks across moss-covered ground. The camera tracks its gentle movement in shallow focus. Natural ambiance fills the soundscape — birds chirping, distant rustle of leaves, and a light breeze passing through the forest.", "duration": 6, "resolution": "720p", "aspect_ratio": "16:9", "person_generation": "allow_all" } ) result = await response.get() print(result["output"]) asyncio.run(main()) ``` ## Additional Resources - [Playground](https://modelrunner.ai/models/google/veo-3.1/text-to-video) - [OpenAPI Schema](https://modelrunner.ai/models/google/veo-3.1/text-to-video/openapi.json) - [LLM Instructions](https://modelrunner.ai/models/google/veo-3.1/text-to-video/llms.txt)