Ana içeriğe geç
meta avatarı

SAM Audio — Separate API

meta/sam-audio/separate

Herhangi bir sesi düz bir dille tarif ederek miks bir kayıttan izole edin.

segmentasyon
üretilen ses için saniye başına 0.001667

Model girdisi

Input

URL of the audio file to process (WAV, MP3, FLAC supported).

Text prompt describing the sound to isolate.

Additional Settings

Customize your input with more control.

Automatically predict temporal spans where the target sound occurs.

Min: 1 - Max: 7

Number of candidates to generate and rank. Values above 1 incur an additional charge per extra candidate.

The acceleration level to use, trading speed against quality.

Min: 10 - Max: 60

Maximum audio duration (seconds) to process in a single pass.

Min: 0 - Max: 30

Overlap duration (seconds) between chunks for crossfade blending.

Output audio format.

You need to be logged in to run this model and view results.
Log in

Model çıktısı

Output

Loading
Generated in 5.706 seconds
Logs (1 lines)

Örnek istekler

Örnekler

SAM Audio — Separate API

SAM Audio — Separate is a audio-to-audio AI model by meta. On ModelRunner it runs through a REST API or via MCP from any AI assistant, at $0.001667 per second of audio.

POST https://queue.modelrunner.run/meta/sam-audio/separate

cURL

# Submit a request to the queue. Input fields go at the top level of the
# body. The optional reserved "metadata" object holds your own flat string
# tags — stored on the request, never sent to the model; filter later with
# GET https://queue.modelrunner.run/requests?metadata=<url-encoded JSON>.
curl -X POST https://queue.modelrunner.run/meta/sam-audio/separate \
  -H "Authorization: Key $MRUN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "dog barking",
    "audio_url": "https://media.modelrunner.ai/JbyO9TzNupSFoqNrlmaWe.mp3",
    "acceleration": "balanced",
    "chunk_overlap": 5,
    "output_format": "wav",
    "predict_spans": false,
    "max_chunk_duration": 60,
    "reranking_candidates": 1,
    "metadata": {
      "project": "my-project"
    }
  }'
# → { "request_id": "...", "status_url": "...", "response_url": "..." }

# Poll status_url until "COMPLETED", then fetch the result
curl "https://queue.modelrunner.run/meta/sam-audio/separate/requests/$REQUEST_ID/status" \
  -H "Authorization: Key $MRUN_API_KEY"
curl "https://queue.modelrunner.run/meta/sam-audio/separate/requests/$REQUEST_ID" \
  -H "Authorization: Key $MRUN_API_KEY"

JavaScript

import { modelrunner } from "@modelrunner/client";

const result = await modelrunner.subscribe("meta/sam-audio/separate", {
  input: {
    "prompt": "dog barking",
    "audio_url": "https://media.modelrunner.ai/JbyO9TzNupSFoqNrlmaWe.mp3",
    "acceleration": "balanced",
    "chunk_overlap": 5,
    "output_format": "wav",
    "predict_spans": false,
    "max_chunk_duration": 60,
    "reranking_candidates": 1
  },
});
console.log(result);

Python

import os
import requests

headers = {"Authorization": f"Key {os.environ['MRUN_API_KEY']}"}

submitted = requests.post(
    "https://queue.modelrunner.run/meta/sam-audio/separate",
    headers=headers,
    json={
      "prompt": "dog barking",
      "audio_url": "https://media.modelrunner.ai/JbyO9TzNupSFoqNrlmaWe.mp3",
      "acceleration": "balanced",
      "chunk_overlap": 5,
      "output_format": "wav",
      "predict_spans": false,
      "max_chunk_duration": 60,
      "reranking_candidates": 1
    },
).json()

# Poll submitted["status_url"] until "COMPLETED", then:
result = requests.get(submitted["response_url"], headers=headers).json()

Input parameters

Input parameters of SAM Audio — Separate
NameTypeRequiredDescription
audio_urlstring (uri)yesURL of the audio file to process (WAV, MP3, FLAC supported).
promptstringyesText prompt describing the sound to isolate.
predict_spansbooleannoAutomatically predict temporal spans where the target sound occurs. Default: false.
reranking_candidatesintegernoNumber of candidates to generate and rank. Values above 1 incur an additional charge per extra candidate. Default: 1.
accelerationenumnoThe acceleration level to use, trading speed against quality. One of: fast, balanced, quality. Default: "balanced".
max_chunk_durationnumbernoMaximum audio duration (seconds) to process in a single pass. Default: 60.
chunk_overlapnumbernoOverlap duration (seconds) between chunks for crossfade blending. Default: 5.
output_formatenumnoOutput audio format. One of: wav, mp3. Default: "wav".

Machine-readable: OpenAPI schema · llms.txt

Use SAM Audio — Separate from Claude & Cursor (MCP)

Point Claude Code, Claude Desktop, Cursor, or any MCP client at the ModelRunner MCP server and SAM Audio — Separate becomes a tool your assistant can call directly — it authorizes via OAuth (no API key in config) and runs this model with the run_model tool using the endpoint meta/sam-audio/separate.

MCP client config (Claude Desktop, Cursor)

{
  "mcpServers": {
    "modelrunner": {
      "command": "npx",
      "args": ["-y", "mcp-remote", "https://mcp.modelrunner.run/mcp"]
    }
  }
}

Claude Code

claude mcp add --transport http modelrunner https://mcp.modelrunner.run/mcp

Then ask your assistant, for example: “Run meta/sam-audio/separate on ModelRunner to generate audio”. MCP setup guide.

Model Detayları

Model Detayları

SAM Audio (Separate), metinle yönlendirilen ses kaynağı ayrıştırması yapar: miks bir kayıt ve istediğiniz sesi anlatan kısa, doğal dilde bir açıklama verirsiniz; model o sesi ayrı bir parça olarak izole eder. Hedefi nasıl söylüyorsanız öyle tarif edin — "köpek havlaması", "ana vokal", "konuşma", "yağmur", "elektro gitar" — model de izole edilmiş hedefi barındırılan bir ses dosyası olarak döndürür. Güçlü yanı open-vocabulary hedeflemedir: sabit bir stem listesiyle sınırlı değilsiniz, bu yüzden insan seslerinde, enstrümanlarda, çevre seslerinde ve efektlerde aynı şekilde çalışır. Model ayrıca geri kalan sesi (hedef dışındaki her şeyi) de hesaplar, ancak bu endpoint sonuç olarak izole edilmiş hedef parçayı döndürür.

## En uygun olduğu işler - Gürültülü ya da katmanlı bir miksten tek bir insan sesini, enstrümanı ya da ses efektini çekip almak - Hazır bir stem yerine sözcüklerle tarif edilen herhangi bir sesi izole etmek (ör. "havlayan bir köpek", "kalabalık alkışı") - Yalnızca ilgilendiğiniz sesi çıkararak saha kayıtlarını, podcast’leri ve röportajları temizlemek - Tek bir öğenin ayrı olarak gerektiği düzenleme, sampling ya da erişilebilirlik işleri için kaydı hazırlamak

## Şu durumlarda başka bir model seçin - Yalnızca arka plandaki gürültüyü ya da müziği temizleyip net bir konuşma elde etmeniz gerekiyorsa — özel bir ses izolasyonu modeli daha basittir - Tarif edilen tek bir hedef yerine sabit bir çoklu stem ayrımının (vokal/davul/bas/diğer) tek seferde döndürülmesini istiyorsanız - Mevcut bir kaydı ayrıştırmak yerine ses üretmeniz ya da dönüştürmeniz gerekiyorsa — bir text-to-audio ya da audio-to-audio sentez modeli kullanın

## İpuçları - `prompt` alanını kısa ve somut tutun; sesi, birine anlatırken kullanacağınız sözcüklerle adlandırın - `audio_url` için WAV, MP3 ya da FLAC dosyası verin - `acceleration` hız ile kalite arasında denge kurar (`fast`, `balanced`, `quality`); varsayılan `balanced` iyi bir başlangıç noktasıdır - `reranking_candidates` kaliteyi artırmak için birden fazla ayrıştırma üretip bunları sıralar — 1’in üzerindeki değerlerde her ek aday için ayrıca ücret alındığını unutmayın - Uzun dosyalarda `max_chunk_duration` ve `chunk_overlap`, sesin geçişler boyunca nasıl bölüneceğini ve crossfade ile nasıl birleştirileceğini belirler

## Sınırlamalar - Yoğun biçimde üst üste binen ya da spektral olarak benzer sesler, hedef ile geri kalan ses arasında birbirine sızabilir - Endpoint yalnızca izole edilmiş hedefi döndürür; geri kalan ses hesaplanır ama sonuç olarak döndürülmez

ModelRunner JavaScript client ile çalıştırmak için: ```js import { modelrunner } from "@modelrunner/client";

const result = await modelrunner.subscribe("meta/sam-audio/separate", { input: { audio_url: "https://media.modelrunner.ai/JbyO9TzNupSFoqNrlmaWe.mp3", prompt: "dog barking", acceleration: "balanced", }, }); ```