Ana içeriğe geç
bytedance avatarı

OmniHuman 1.5 API

bytedance/omnihuman/v1.5

Bir kişinin durağan fotoğrafını ses kaydıyla senkronize biçimde konuşturup hareket ettirin; doğal bir talking-head videosu elde edin.

üretilen video için saniye başına 0.16

Model girdisi

Input

URL of the source photo of the person to animate.

URL of the audio track the person should speak or sing. Keep audio under 30s at 1080p, under 60s at 720p.

Optional text prompt guiding the motion, gestures, and performance. Leave empty to let the audio drive the animation.

Optional mask image. When the photo has more than one person, only the person inside the white region of the mask will be animated to speak.

Generate faster with a slight quality trade-off. No price impact.

Output resolution. 1080p limits input audio to 30s; 720p allows up to 60s.

You need to be logged in to run this model and view results.
Log in

Model çıktısı

Output

Loading
Generated in 169.322 seconds
Logs (1 lines)

Örnek istekler

Örnekler

Example output 1Example output 2

OmniHuman 1.5 API

OmniHuman 1.5 is a image-to-video AI model by bytedance. On ModelRunner it runs through a REST API or via MCP from any AI assistant, at $0.16 per second of video.

POST https://queue.modelrunner.run/bytedance/omnihuman/v1.5

cURL

# Submit a request to the queue. Input fields go at the top level of the
# body. The optional reserved "metadata" object holds your own flat string
# tags — stored on the request, never sent to the model; filter later with
# GET https://queue.modelrunner.run/requests?metadata=<url-encoded JSON>.
curl -X POST https://queue.modelrunner.run/bytedance/omnihuman/v1.5 \
  -H "Authorization: Key $MRUN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "natural presenter delivering the line, subtle head movement",
    "audio_url": "https://media.modelrunner.ai/v293WP0BkvLcoXC3MJLud.mp3",
    "image_url": "https://media.modelrunner.ai/ho4HXHjCrHjv7MZs-omnihuman_v15_input_image.png",
    "resolution": "720p",
    "turbo_mode": true,
    "metadata": {
      "project": "my-project"
    }
  }'
# → { "request_id": "...", "status_url": "...", "response_url": "..." }

# Poll status_url until "COMPLETED", then fetch the result
curl "https://queue.modelrunner.run/bytedance/omnihuman/v1.5/requests/$REQUEST_ID/status" \
  -H "Authorization: Key $MRUN_API_KEY"
curl "https://queue.modelrunner.run/bytedance/omnihuman/v1.5/requests/$REQUEST_ID" \
  -H "Authorization: Key $MRUN_API_KEY"

JavaScript

import { modelrunner } from "@modelrunner/client";

const result = await modelrunner.subscribe("bytedance/omnihuman/v1.5", {
  input: {
    "prompt": "natural presenter delivering the line, subtle head movement",
    "audio_url": "https://media.modelrunner.ai/v293WP0BkvLcoXC3MJLud.mp3",
    "image_url": "https://media.modelrunner.ai/ho4HXHjCrHjv7MZs-omnihuman_v15_input_image.png",
    "resolution": "720p",
    "turbo_mode": true
  },
});
console.log(result);

Python

import os
import requests

headers = {"Authorization": f"Key {os.environ['MRUN_API_KEY']}"}

submitted = requests.post(
    "https://queue.modelrunner.run/bytedance/omnihuman/v1.5",
    headers=headers,
    json={
      "prompt": "natural presenter delivering the line, subtle head movement",
      "audio_url": "https://media.modelrunner.ai/v293WP0BkvLcoXC3MJLud.mp3",
      "image_url": "https://media.modelrunner.ai/ho4HXHjCrHjv7MZs-omnihuman_v15_input_image.png",
      "resolution": "720p",
      "turbo_mode": true
    },
).json()

# Poll submitted["status_url"] until "COMPLETED", then:
result = requests.get(submitted["response_url"], headers=headers).json()

Input parameters

Input parameters of OmniHuman 1.5
NameTypeRequiredDescription
image_urlstring (uri)yesURL of the source photo of the person to animate.
audio_urlstring (uri)yesURL of the audio track the person should speak or sing. Keep audio under 30s at 1080p, under 60s at 720p.
promptstringnoOptional text prompt guiding the motion, gestures, and performance. Leave empty to let the audio drive the animation.
mask_urlstringnoOptional mask image. When the photo has more than one person, only the person inside the white region of the mask will be animated to speak.
turbo_modebooleannoGenerate faster with a slight quality trade-off. No price impact. Default: false.
resolutionenumnoOutput resolution. 1080p limits input audio to 30s; 720p allows up to 60s. One of: 720p, 1080p. Default: "1080p".

Machine-readable: OpenAPI schema · llms.txt

Use OmniHuman 1.5 from Claude & Cursor (MCP)

Point Claude Code, Claude Desktop, Cursor, or any MCP client at the ModelRunner MCP server and OmniHuman 1.5 becomes a tool your assistant can call directly — it authorizes via OAuth (no API key in config) and runs this model with the run_model tool using the endpoint bytedance/omnihuman/v1.5.

MCP client config (Claude Desktop, Cursor)

{
  "mcpServers": {
    "modelrunner": {
      "command": "npx",
      "args": ["-y", "mcp-remote", "https://mcp.modelrunner.run/mcp"]
    }
  }
}

Claude Code

claude mcp add --transport http modelrunner https://mcp.modelrunner.run/mcp

Then ask your assistant, for example: “Run bytedance/omnihuman/v1.5 on ModelRunner to generate video”. MCP setup guide.

Model Detayları

Model Detayları

OmniHuman 1.5, bir kişinin tek bir durağan fotoğrafını ve bir ses kaydını talking-head videosuna dönüştürür: öznenin ağzı, mimikleri ve baş/vücut hareketleri sesteki konuşma ya da şarkıya uyacak şekilde canlandırılır; genel hareketi ve performansı yönlendirmek için isteğe bağlı bir metin prompt’u verebilirsiniz. Net bir portre ve söylemesini istediğiniz sesi verin, o kişinin bunu söylediği bir MP4 alın. Güçlü yanı, tek bir görselden sıkı ses güdümlü dudak ve yüz senkronu ile inandırıcı, robotik olmayan vücut hareketi üretmesidir.

## En uygun olduğu işler - Bir vesikalık ya da portreyi, seslendirme veya anlatıma dudak senkronu yapan bir sözcü videosuna dönüştürmek - Açıklayıcı videolar, reklamlar, eğitimler ve ürün demoları için avatar ve dijital sunucu klipleri - Bir karakteri ya da illüstrasyonu verilen ses kaydıyla eş zamanlı şarkı söyletmek veya konuşturmak - Çevrilmiş sesin hazır olduğu yerelleştirilmiş/dublajlı talking-head klipleri - Tek fotoğraf ve sesten hızlı sosyal medya ya da UGC tarzı talking-head içerik

## Şu durumlarda başka bir model seçin - Elinizde zaten bir talking-head video varsa ve durağan fotoğraf canlandırmak yerine yalnızca dudakları yeni sese göre yeniden senkronlamak istiyorsanız — bir lip-sync (video-to-video) modeli kullanın - Ses olmadan, yalnızca metin prompt’undan bir sahneyi ya da nesneyi canlandırmak istiyorsanız — image-to-video ya da text-to-video modeli kullanın - Hareketi yönlendirecek konuşma ve ses kaydı olmayan sıradan bir image-to-video klip gerekiyorsa — standart bir image-to-video modeli kullanın

## İpuçları - En doğru senkron için yüzün tamamının görünür ve engelsiz olduğu, net, önden çekilmiş bir fotoğraf kullanın - Ses süresi çözünürlüğe göre sınırlıdır: 1080p’de sesi 30s’nin, 720p’de 60s’nin altında tutun - Jestleri, enerjiyi ve kamera hissini yönlendirmek için `prompt` kullanın (ör. "sakin sunucu, hafif baş sallama"); her şeyi sesin yönetmesi için boş bırakın - Fotoğrafta birden fazla kişi varsa `mask_url` ile bir maske görseli verin — yalnızca maskenin beyaz bölgesindeki kişi konuşturulur - `turbo_mode`, ek ücret olmadan daha hızlı üretim karşılığında kaliteden biraz ödün verir

## Sınırlamalar - Ses kaydıyla yönlendirilen insan özneler için tasarlandı; insan olmayan özneler ya da net konuşma/vokal sinyali olmayan sesler daha zayıf sonuç verir - Çok uzun klipler, çözünürlük başına ses süresi sınırlarına uymak için bölünmelidir

ModelRunner JavaScript client ile çalıştırmak için: ```js import { modelrunner } from "@modelrunner/client";

const result = await modelrunner.subscribe("bytedance/omnihuman/v1.5", { input: { image_url: "https://media.modelrunner.ai/ho4HXHjCrHjv7MZs-omnihuman_v15_input_image.png", audio_url: "https://media.modelrunner.ai/v293WP0BkvLcoXC3MJLud.mp3", prompt: "natural presenter delivering the line, subtle head movement", resolution: "1080p", }, }); ```