Ana içeriğe geç
bytedance avatarı

LatentSync 1.0 API

bytedance/latentsync

Bir talking-head videosundaki ağzı, sese koşullu diffusion ile yeni bir ses kaydına göre yeniden senkronlayın — çıktı saniyesi başına $0.014 ile katalogdaki en düşük fiyatlı lip-sync.

düzenleme
üretilen video için saniye başına 0.014

Model girdisi

Input

URL of the source video — an MP4 with one clearly visible, front-facing speaker in every frame; the run fails on any frame where no face is detected. The result keeps this video's frame size, and its identity, lighting and background; only the mouth region is regenerated. Video running past the end of the audio is discarded, so trim it to roughly the audio's length.

URL of the speech track the speaker should appear to say (MP3, AAC, WAV or M4A). The result runs for the shorter of this track and the video, trimmed down to a whole multiple of 0.64 seconds — supply audio slightly shorter than the video and expect the last fraction of a second to be cut.

Additional Settings

Customize your input with more control.

Min: 0 - Max: 10

Strength of the audio conditioning during diffusion. The default of 1 leaves classifier-free guidance off; values above 1 switch it on. The model's own demo exposes 1–3.5, though the field accepts up to 10.

Random seed. 0 (the default) draws a fresh random seed on every run, so repeated calls with identical inputs differ; any positive integer is used as given for a repeatable run.

You need to be logged in to run this model and view results.
Log in

Model çıktısı

Output

Loading
Generated in 43.244 seconds
Logs (1 lines)

Örnek istekler

Örnekler

Example output 1Example output 2

LatentSync 1.0 API

LatentSync 1.0 is a video-to-video AI model by bytedance. On ModelRunner it runs through a REST API or via MCP from any AI assistant, at $0.014 per second of video.

POST https://queue.modelrunner.run/bytedance/latentsync

cURL

# Submit a request to the queue. Input fields go at the top level of the
# body. The optional reserved "metadata" object holds your own flat string
# tags — stored on the request, never sent to the model; filter later with
# GET https://queue.modelrunner.run/requests?metadata=<url-encoded JSON>.
curl -X POST https://queue.modelrunner.run/bytedance/latentsync \
  -H "Authorization: Key $MRUN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "seed": 0,
    "audio_url": "https://media.modelrunner.ai/r2Nq5Eh8szuYnxfAVRVDG.wav",
    "video_url": "https://media.modelrunner.ai/Waq0PrjekwA7pCyykfy0Z.mp4",
    "guidance_scale": 1,
    "metadata": {
      "project": "my-project"
    }
  }'
# → { "request_id": "...", "status_url": "...", "response_url": "..." }

# Poll status_url until "COMPLETED", then fetch the result
curl "https://queue.modelrunner.run/bytedance/latentsync/requests/$REQUEST_ID/status" \
  -H "Authorization: Key $MRUN_API_KEY"
curl "https://queue.modelrunner.run/bytedance/latentsync/requests/$REQUEST_ID" \
  -H "Authorization: Key $MRUN_API_KEY"

JavaScript

import { modelrunner } from "@modelrunner/client";

const result = await modelrunner.subscribe("bytedance/latentsync", {
  input: {
    "seed": 0,
    "audio_url": "https://media.modelrunner.ai/r2Nq5Eh8szuYnxfAVRVDG.wav",
    "video_url": "https://media.modelrunner.ai/Waq0PrjekwA7pCyykfy0Z.mp4",
    "guidance_scale": 1
  },
});
console.log(result);

Python

import os
import requests

headers = {"Authorization": f"Key {os.environ['MRUN_API_KEY']}"}

submitted = requests.post(
    "https://queue.modelrunner.run/bytedance/latentsync",
    headers=headers,
    json={
      "seed": 0,
      "audio_url": "https://media.modelrunner.ai/r2Nq5Eh8szuYnxfAVRVDG.wav",
      "video_url": "https://media.modelrunner.ai/Waq0PrjekwA7pCyykfy0Z.mp4",
      "guidance_scale": 1
    },
).json()

# Poll submitted["status_url"] until "COMPLETED", then:
result = requests.get(submitted["response_url"], headers=headers).json()

Input parameters

Input parameters of LatentSync 1.0
NameTypeRequiredDescription
video_urlstring (uri)yesURL of the source video — an MP4 with one clearly visible, front-facing speaker in every frame; the run fails on any frame where no face is detected. The result keeps this video's frame size, and its identity, lighting and background; only the mouth region is regenerated. Video running past the end of the audio is discarded, so trim it to roughly the audio's length.
audio_urlstring (uri)yesURL of the speech track the speaker should appear to say (MP3, AAC, WAV or M4A). The result runs for the shorter of this track and the video, trimmed down to a whole multiple of 0.64 seconds — supply audio slightly shorter than the video and expect the last fraction of a second to be cut.
guidance_scalenumbernoStrength of the audio conditioning during diffusion. The default of 1 leaves classifier-free guidance off; values above 1 switch it on. The model's own demo exposes 1–3.5, though the field accepts up to 10. Default: 1.
seedintegernoRandom seed. 0 (the default) draws a fresh random seed on every run, so repeated calls with identical inputs differ; any positive integer is used as given for a repeatable run. Default: 0.

Machine-readable: OpenAPI schema · llms.txt

Use LatentSync 1.0 from Claude & Cursor (MCP)

Point Claude Code, Claude Desktop, Cursor, or any MCP client at the ModelRunner MCP server and LatentSync 1.0 becomes a tool your assistant can call directly — it authorizes via OAuth (no API key in config) and runs this model with the run_model tool using the endpoint bytedance/latentsync.

MCP client config (Claude Desktop, Cursor)

{
  "mcpServers": {
    "modelrunner": {
      "command": "npx",
      "args": ["-y", "mcp-remote", "https://mcp.modelrunner.run/mcp"]
    }
  }
}

Claude Code

claude mcp add --transport http modelrunner https://mcp.modelrunner.run/mcp

Then ask your assistant, for example: “Run bytedance/latentsync on ModelRunner to generate video”. MCP setup guide.

Model Detayları

Model Detayları

LatentSync 1.0, bir talking-head videosunu ve ayrı bir konuşma kaydını alır, konuşmacının ağzını yeni sese uyacak şekilde yeniden canlandırır. Kimlik, ışık ve arka plan kaynaktan aynen gelir; yalnızca ağız yeniden üretilir ve sonuç girdinin kare boyutunu korur. Yöntem sese koşullu latent diffusion’dır: Whisper ses embedding’leri, arada landmark ya da hareket temsili olmadan cross-attention ile diffusion modeline aktarılır; TREPA, LPIPS ve SyncNet loss fonksiyonları da ağız hareketini kareden kareye kararlı tutar. Teslim edilen video saniyesi başına $0.014 ile — 20 saniyelik bir klip için yaklaşık $0.28 — katalogdaki en düşük fiyatlı lip-sync modelidir. Net görünen, önden çekilmiş tek bir konuşmacının olduğu bir MP4 ile MP3, AAC, WAV ya da M4A formatında bir konuşma kaydı verin; metin prompt’u yoktur.

## En uygun olduğu işler - Talking-head görüntülerini, dudaklar çevrilmiş seslendirmeye uysun diye başka bir dile dublajlamak - Bir röportajı, eğitim videosunu ya da reklamı daha temiz ya da yeniden kaydedilmiş bir ses çekimiyle yeniden seslendirmek - Kaydedilen konuşma ile ağız hareketlerinin senkrondan kaydığı bir çekimi düzeltmek - Bir text-to-speech modelinin ardına eklemek: önce sesi üretin, ardından görüntüyü bu sese göre lip-sync edin - Saniye başına maliyetin önemli olduğu yüksek hacimli ya da uzun soluklu lip-sync işleri

## Şu durumlarda başka bir model seçin - Süre uyumsuzluğunun nasıl çözüleceğini kontrol etmeniz gerekiyorsa — bu model her zaman ikisinden kısa olana göre keser; `sync/lipsync/v2` ise loop, bounce, silence ve remap seçenekleriyle `sync_mode` sunar - Elinizde görüntü değil durağan bir fotoğraf varsa — bir portreyi canlandırmak için `bytedance/omnihuman/v1.5` ya da `wan-video/wan/v2.7/image-to-video/audio-driven` kullanın - Konuşmayı henüz üretmediyseniz — bu endpoint hazır bir ses dosyası alır ve metin prompt’u almaz; önce bir text-to-speech modeli çalıştırıp çıktısını buraya verin

## İpuçları - Sesi videodan biraz kısa verin: sonuç ikisinden kısa olanın süresi kadar sürer ve 0.64 saniyenin tam katına kırpılır, bu yüzden konuşmanın sonda kalan küçük bir parçası düşer - Kaynak videoyu kabaca ses uzunluğuna kırpın; fazla video atılır ve yalnızca çalıştırmayı yavaşlatır

## Sınırlamalar - Her karede algılanabilir bir yüz olmalı: ara görüntü ya da yüz içermeyen bir kare tüm çalıştırmayı başarısız kılar (ücret alınmaz) - Kare hızı ayarlanamaz — sonuç 25 fps olarak teslim edilir, diğer kare hızları yeniden örneklenir - Ağız detayı sınırlıdır; bu nesil yüksek çözünürlükte eğitilmedi - Ses bandı 16 kHz mono AAC olarak yeniden kodlanıp döner

ModelRunner JavaScript client ile çalıştırmak için: ```js import { modelrunner } from "@modelrunner/client";

const result = await modelrunner.subscribe("bytedance/latentsync", { input: { video_url: "https://media.modelrunner.ai/NLaN6i1oQvbd8Mn4oEDXs.mp4", audio_url: "https://media.modelrunner.ai/lZfxwe6ZN6ZikXQhVtiq7.wav", }, }); ```