# Wan 2.7 Image to Video (Audio Driven) > Drive a still photo with your own audio clip: the track is used for lip-sync and action timing, producing a 2-15 second video at 720P or 1080P that performs in time with the sound. ## Overview - **Endpoint**: `https://queue.modelrunner.run/wan-video/wan/v2.7/image-to-video/audio-driven` - **Model ID**: `wan-video/wan/v2.7/image-to-video/audio-driven` - **Category**: image-to-video - **Kind**: inference - **Tags**: wan, wan-2.7, image-to-video, audio-driven, lip-sync, lipsync, talking-photo, talking-head, avatar, video-generation, animate-photo ## Pricing - **720P**: $0.1 per output second - **1080P**: $0.15 per output second ## Request Lifecycle This model runs on the ModelRunner **asynchronous queue API** — a single POST does not return the output. Every call requires an `Authorization: Key $MODEL_RUNNER_KEY` header. Run three steps: 1. **Submit** — `POST https://queue.modelrunner.run/wan-video/wan/v2.7/image-to-video/audio-driven` with a JSON body holding the input fields at the top level. The body may also include a reserved top-level `metadata` object — a flat string map (max 16 keys, key ≤64 / value ≤512 chars) stored on the request for your own tagging. It is never sent to the model; filter your request history with `GET https://queue.modelrunner.run/requests?metadata=` (exact key=value matches, AND-ed). The response carries request handles only (no output yet): ```json { "status": "IN_QUEUE", "request_id": "<21-char id>", "status_url": "https://queue.modelrunner.run/wan-video/wan/v2.7/image-to-video/audio-driven/requests//status", "response_url": "https://queue.modelrunner.run/wan-video/wan/v2.7/image-to-video/audio-driven/requests/", "cancel_url": "https://queue.modelrunner.run/wan-video/wan/v2.7/image-to-video/audio-driven/requests//cancel" } ``` 2. **Poll status** — `GET ` until `status` is `COMPLETED`. Possible values are `IN_QUEUE`, `IN_PROGRESS`, `COMPLETED`, `FAILED`, `CANCELLED`. A `FAILED` request responds with HTTP 400 and an `error` field. 3. **Read result** — `GET `. Returns the finished request, including the generated `output`: ```json { "id": "", "status": "COMPLETED", "output": ..., "input": ... } ``` The JavaScript and Python SDKs below perform steps 2–3 for you. In any language without an SDK (Swift, Go, Kotlin, etc.) you must implement the polling loop and the final result fetch yourself — see the cURL example for the full flow. ### Input Schema - **`seed`** (`integer`, _optional_): Random seed for reproducible results. Omit for a random seed each run. - Range: `0` to `2147483647` - **`prompt`** (`string`, _optional_): Optional description of the motion and camera movement. The scene is already fixed by the start image and the mouth movement and action timing come from the driving audio, so use the prompt for gesture, framing and camera. Chinese and English are supported. - **`duration`** (`integer`, _optional_): Length of the generated video in seconds (2-15). This is the output length, and the only length that is billed - the driving audio clip's own length never changes it. - Default: `5` - Range: `2` to `15` - **`resolution`** (`ResolutionEnum`, _optional_): Output video resolution. 720P bills at $0.10 per second of finished video; 1080P (default) bills at $0.15 per second. - Default: `"1080P"` - Options: `"720P"`, `"1080P"` - **`end_image_url`** (`string`, _optional_): Optional closing frame. Supply it to pin where the clip ends while the driving audio times everything in between; it cannot be used on its own, without a start frame. Same formats and size limits as the start frame, and it should share the start frame's aspect ratio. - **`negative_prompt`** (`string`, _optional_): Describe content to avoid in the generated video. - **`start_image_url`** (`string`, _required_): The opening frame the video animates from. For lip-sync, pick a frame where the subject's face and mouth are clearly visible and unobscured. JPEG, JPG, PNG (alpha channel not supported), BMP or WEBP; width and height each between 240 and 8000 px, aspect ratio between 1:8 and 8:1, up to 20 MB. The finished clip takes its frame shape from this image. - **`driving_audio_url`** (`string`, _required_): The audio clip that drives the performance: the model uses it as the source for lip-sync and action timing, and it is the sound heard in the finished clip. WAV or MP3, 2-30 seconds, up to 15 MB. Audio longer than the requested duration is truncated to the first duration seconds; audio shorter than the requested duration leaves the rest of the clip silent, so match the two for sound throughout. - **`enable_prompt_expansion`** (`boolean`, _optional_): When enabled, an LLM rewrites and enriches your prompt before generation. Disable to follow your exact wording. - Default: `true` ### Output Schema _No `Output` schema properties are available._ ## Default Example **Input** ```json { "prompt": "The ceramicist speaks warmly to the camera, her hands moving gently as she talks, dust motes drifting in the window light", "duration": 5, "resolution": "720P", "start_image_url": "https://media.modelrunner.ai/wCzkJnfwOMO3G6fEXkPCg.png", "driving_audio_url": "https://media.modelrunner.ai/N1Mq2Pphb4vCQkhIWg8HZ.mp3", "enable_prompt_expansion": true } ``` **Output** ```json "https://media.modelrunner.ai/at69YH6ZNKMQnxoFjb830.mp4" ``` ## Usage Examples ### cURL The queue API is asynchronous: submit the request, poll `status_url` until it is `COMPLETED`, then read the result from `response_url`. Requires `jq`. ```bash # 1. Submit the request (returns request handles, not the output) SUBMIT=$(curl --silent --request POST \ --url https://queue.modelrunner.run/wan-video/wan/v2.7/image-to-video/audio-driven \ --header "Authorization: Key $MODEL_RUNNER_KEY" \ --header "Content-Type: application/json" \ --data '{ "prompt": "The ceramicist speaks warmly to the camera, her hands moving gently as she talks, dust motes drifting in the window light", "duration": 5, "resolution": "720P", "start_image_url": "https://media.modelrunner.ai/wCzkJnfwOMO3G6fEXkPCg.png", "driving_audio_url": "https://media.modelrunner.ai/N1Mq2Pphb4vCQkhIWg8HZ.mp3", "enable_prompt_expansion": true }') STATUS_URL=$(echo "$SUBMIT" | jq -r '.status_url') RESPONSE_URL=$(echo "$SUBMIT" | jq -r '.response_url') # 2. Poll until the request leaves the queue / in-progress state while true; do STATUS=$(curl --silent --url "$STATUS_URL" \ --header "Authorization: Key $MODEL_RUNNER_KEY" | jq -r '.status') echo "Status: $STATUS" case "$STATUS" in COMPLETED) break ;; FAILED|CANCELLED) echo "Request $STATUS"; exit 1 ;; esac sleep 1 done # 3. Read the finished request, including the generated output curl --silent --url "$RESPONSE_URL" \ --header "Authorization: Key $MODEL_RUNNER_KEY" ``` ### JavaScript ```javascript import { modelrunner } from "@modelrunner/client"; const result = await modelrunner.subscribe("wan-video/wan/v2.7/image-to-video/audio-driven", { input: { "prompt": "The ceramicist speaks warmly to the camera, her hands moving gently as she talks, dust motes drifting in the window light", "duration": 5, "resolution": "720P", "start_image_url": "https://media.modelrunner.ai/wCzkJnfwOMO3G6fEXkPCg.png", "driving_audio_url": "https://media.modelrunner.ai/N1Mq2Pphb4vCQkhIWg8HZ.mp3", "enable_prompt_expansion": true } }); console.log(result.data); ``` ### Python ```python import asyncio import modelrunner_ai async def main(): response = await modelrunner_ai.submit_async( "wan-video/wan/v2.7/image-to-video/audio-driven", arguments={ "prompt": "The ceramicist speaks warmly to the camera, her hands moving gently as she talks, dust motes drifting in the window light", "duration": 5, "resolution": "720P", "start_image_url": "https://media.modelrunner.ai/wCzkJnfwOMO3G6fEXkPCg.png", "driving_audio_url": "https://media.modelrunner.ai/N1Mq2Pphb4vCQkhIWg8HZ.mp3", "enable_prompt_expansion": true } ) result = await response.get() print(result["output"]) asyncio.run(main()) ``` ## Additional Resources - [Playground](https://modelrunner.ai/models/wan-video/wan/v2.7/image-to-video/audio-driven) - [OpenAPI Schema](https://modelrunner.ai/models/wan-video/wan/v2.7/image-to-video/audio-driven/openapi.json) - [LLM Instructions](https://modelrunner.ai/models/wan-video/wan/v2.7/image-to-video/audio-driven/llms.txt)