Ana içeriğe geç
elevenlabs avatarı

ElevenLabs Scribe v1 API

elevenlabs/scribe-v1

Konuşma kayıtlarını 99 dilde, kelime düzeyinde zaman damgaları, konuşmacı etiketleri ve ses olayı etiketleriyle birlikte doğru metne dönüştürün.

0.03

Model girdisi

Input

URL of the audio file to transcribe. Supported formats: mp3, wav, m4a, ogg, aac.

Additional Settings

Customize your input with more control.

ISO-639 language code of the audio (e.g. eng, spa, fra, deu, jpn). Leave unset to auto-detect the spoken language.

Tag non-speech audio events like laughter and applause inline in the transcript.

Annotate which speaker said each word.

You need to be logged in to run this model and view results.
Log in

Model çıktısı

Output

Welcome to the show. Today we're talking about how AI assistants actually run models. Let's start simple. What is an MCP server? Thanks for having me. An MCP server is a bridge. Your assistant connects to it, discovers a set of tools, and can call them mid-conversation. In our case, those tools run image, video, and audio models in the cloud. And what happens to the files? Say I record an interview like this one. You upload it once, it becomes a hosted URL, and any transcription model can take it from there. The transcript comes back into the conversation with speaker turns.

Generated in 3.004 seconds
Logs (1 lines)

Örnek istekler

Örnekler

ElevenLabs Scribe v1 API

ElevenLabs Scribe v1 is a speech-to-text AI model by elevenlabs. On ModelRunner it runs through a REST API or via MCP from any AI assistant, at $0.03 per response.

POST https://queue.modelrunner.run/elevenlabs/scribe-v1

cURL

# Submit a request to the queue. Input fields go at the top level of the
# body. The optional reserved "metadata" object holds your own flat string
# tags — stored on the request, never sent to the model; filter later with
# GET https://queue.modelrunner.run/requests?metadata=<url-encoded JSON>.
curl -X POST https://queue.modelrunner.run/elevenlabs/scribe-v1 \
  -H "Authorization: Key $MRUN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "diarize": true,
    "audio_url": "https://media.modelrunner.ai/U74wCnE5oqjyctCo-interview.m4a",
    "tag_audio_events": true,
    "metadata": {
      "project": "my-project"
    }
  }'
# → { "request_id": "...", "status_url": "...", "response_url": "..." }

# Poll status_url until "COMPLETED", then fetch the result
curl "https://queue.modelrunner.run/elevenlabs/scribe-v1/requests/$REQUEST_ID/status" \
  -H "Authorization: Key $MRUN_API_KEY"
curl "https://queue.modelrunner.run/elevenlabs/scribe-v1/requests/$REQUEST_ID" \
  -H "Authorization: Key $MRUN_API_KEY"

JavaScript

import { modelrunner } from "@modelrunner/client";

const result = await modelrunner.subscribe("elevenlabs/scribe-v1", {
  input: {
    "diarize": true,
    "audio_url": "https://media.modelrunner.ai/U74wCnE5oqjyctCo-interview.m4a",
    "tag_audio_events": true
  },
});
console.log(result);

Python

import os
import requests

headers = {"Authorization": f"Key {os.environ['MRUN_API_KEY']}"}

submitted = requests.post(
    "https://queue.modelrunner.run/elevenlabs/scribe-v1",
    headers=headers,
    json={
      "diarize": true,
      "audio_url": "https://media.modelrunner.ai/U74wCnE5oqjyctCo-interview.m4a",
      "tag_audio_events": true
    },
).json()

# Poll submitted["status_url"] until "COMPLETED", then:
result = requests.get(submitted["response_url"], headers=headers).json()

Input parameters

Input parameters of ElevenLabs Scribe v1
NameTypeRequiredDescription
audio_urlstring (uri)yesURL of the audio file to transcribe. Supported formats: mp3, wav, m4a, ogg, aac.
language_codestringnoISO-639 language code of the audio (e.g. eng, spa, fra, deu, jpn). Leave unset to auto-detect the spoken language.
tag_audio_eventsbooleannoTag non-speech audio events like laughter and applause inline in the transcript. Default: true.
diarizebooleannoAnnotate which speaker said each word. Default: true.

Machine-readable: OpenAPI schema · llms.txt

Use ElevenLabs Scribe v1 from Claude & Cursor (MCP)

Point Claude Code, Claude Desktop, Cursor, or any MCP client at the ModelRunner MCP server and ElevenLabs Scribe v1 becomes a tool your assistant can call directly — it authorizes via OAuth (no API key in config) and runs this model with the run_model tool using the endpoint elevenlabs/scribe-v1.

MCP client config (Claude Desktop, Cursor)

{
  "mcpServers": {
    "modelrunner": {
      "command": "npx",
      "args": ["-y", "mcp-remote", "https://mcp.modelrunner.run/mcp"]
    }
  }
}

Claude Code

claude mcp add --transport http modelrunner https://mcp.modelrunner.run/mcp

Then ask your assistant, for example: “Run elevenlabs/scribe-v1 on ModelRunner to generate text”. MCP setup guide.

Model Detayları

Model Detayları

ElevenLabs Scribe v1, konuşma içeren ses dosyalarını doğru bir yazılı metne dönüştürür. Bir ses kaydının URL’sini verin (mp3, wav, m4a, ogg ya da aac); model tam transkripti düz bir string olarak döndürür; doğruluğu 99 dilde sınıfının en iyisidir. Konuşulan dili otomatik algılar, kimin konuştuğunu etiketleyebilir (diarization) ve kahkaha, alkış gibi konuşma dışı ses olaylarını işaretler — bu da onu kayıtları kullanılabilir, aranabilir metne dönüştürmek için güçlü bir varsayılan yapar.

## En uygun olduğu işler - Toplantıları, röportajları, podcast’leri ve sesli notları yazıya dökmek - Kaynak sesten güvenilir kelime sınırlarıyla altyazı ve kapalı altyazı üretmek - Konuşulan dilin bilinmediği ya da karışık olduğu çok dilli transkripsiyon (99 dil, otomatik algılanır) - Diarization ile çok kişili konuşmaların konuşmacıya atfedilmiş transkriptleri - Konuşma kayıtlarından aranabilir arşivler ya da sonrasında çalışacak NLP pipeline’ları kurmak

## Şu durumlarda başka bir model seçin - Yazıya dökmek yerine metinden konuşma üretmek istiyorsanız — bir text-to-speech modeli kullanın - Sesi başka bir dilin metnine çevirmeniz gerekiyorsa — bu model konuşulan dilde yazıya döker, bir konuşma çevirmeni değildir - Devam eden bir görüşmenin canlı, streaming transkripsiyonu gerekiyorsa — bu model yüklenmiş tam bir dosyayı işler ve bitmiş bir transkript döndürür

## İpuçları - Konuşulan dilin otomatik algılanması için `language_code` alanını boş bırakın; dili zaten biliyorsanız ve algılamayı atlamak istiyorsanız bir ISO-639 kodu verin (ör. `eng`, `spa`, `fra`, `deu`, `jpn`). - Çok konuşmacılı kayıtlarda `diarize` alanını etkin (varsayılan) bırakın; model her kelimeyi bir konuşmacıya atar. Tek konuşmacılı seslerde konuşmacı etiketlemesini atlamak için `false` yapın. - Konuşma dışı sesleri (kahkaha, alkış) metnin içinde işaretlemek için `tag_audio_events` alanını etkin (varsayılan) bırakın; yalnızca konuşmadan oluşan temiz bir transkript için `false` yapın. - Net ve yeterince yüksek sesli kaynak kayıtlar kullanın — yoğun arka plan gürültüsü ve üst üste binen konuşmalar doğruluğu düşürür.

## Gelişmiş Yapılandırma - `language_code` (varsayılan: otomatik algılama): transkripsiyon dilini algılamak yerine sabitleyen bir ISO-639 dil kodu. Ses kısa olduğunda ya da dil önceden biliniyorsa işe yarar. - `tag_audio_events` (boolean, varsayılan `true`): `true` olduğunda kahkaha ve alkış gibi konuşma dışı olaylar transkriptin içinde etiketlenir. - `diarize` (boolean, varsayılan `true`): `true` olduğunda her kelimeyi hangi konuşmacının söylediğini işaretler.

ModelRunner JavaScript client ile çalıştırmak için: ```js import { modelrunner } from "@modelrunner/client";

const result = await modelrunner.subscribe("elevenlabs/scribe-v1", { input: { audio_url: "https://storage.googleapis.com/falserverless/web-examples/elevenlabs/sample.mp3", diarize: true, tag_audio_events: true, }, }); ```