Ana içeriğe geç
stability-ai avatarı

Stable Audio 2.5 API

stability-ai/stable-audio-2.5/text-to-audio

Metin prompt’undan uzun soluklu müzik ve ses efektleri üretin — tek çağrıda ~190 saniyeye kadar WAV ses.

0.2

Model girdisi

Input

The prompt to generate audio from. Describe genre, instrumentation, mood, and tempo for music, or the source, environment, and materials for sound effects.

Min: 1 - Max: 190

Duration of the generated audio in seconds (1-190). Billing is a flat rate per generation regardless of length.

Additional Settings

Customize your input with more control.

Min: 4 - Max: 8

Number of denoising steps. More steps can improve quality at the cost of speed.

Min: 1 - Max: 25

Classifier-free guidance scale; higher values follow the prompt more strictly.

Random seed for reproducible generation. Leave empty for a random result.

You need to be logged in to run this model and view results.
Log in

Model çıktısı

Output

Loading
Generated in 3.704 seconds
Logs (1 lines)

Örnek istekler

Örnekler

Stable Audio 2.5 API

Stable Audio 2.5 is a music AI model by stability-ai. On ModelRunner it runs through a REST API or via MCP from any AI assistant, at $0.2 per audio clip.

POST https://queue.modelrunner.run/stability-ai/stable-audio-2.5/text-to-audio

cURL

# Submit a request to the queue. Input fields go at the top level of the
# body. The optional reserved "metadata" object holds your own flat string
# tags — stored on the request, never sent to the model; filter later with
# GET https://queue.modelrunner.run/requests?metadata=<url-encoded JSON>.
curl -X POST https://queue.modelrunner.run/stability-ai/stable-audio-2.5/text-to-audio \
  -H "Authorization: Key $MRUN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "gentle rain on a tin roof with distant rolling thunder and a soft wind, natural field-recording ambience",
    "seconds_total": 20,
    "guidance_scale": 1,
    "num_inference_steps": 8,
    "metadata": {
      "project": "my-project"
    }
  }'
# → { "request_id": "...", "status_url": "...", "response_url": "..." }

# Poll status_url until "COMPLETED", then fetch the result
curl "https://queue.modelrunner.run/stability-ai/stable-audio-2.5/text-to-audio/requests/$REQUEST_ID/status" \
  -H "Authorization: Key $MRUN_API_KEY"
curl "https://queue.modelrunner.run/stability-ai/stable-audio-2.5/text-to-audio/requests/$REQUEST_ID" \
  -H "Authorization: Key $MRUN_API_KEY"

JavaScript

import { modelrunner } from "@modelrunner/client";

const result = await modelrunner.subscribe("stability-ai/stable-audio-2.5/text-to-audio", {
  input: {
    "prompt": "gentle rain on a tin roof with distant rolling thunder and a soft wind, natural field-recording ambience",
    "seconds_total": 20,
    "guidance_scale": 1,
    "num_inference_steps": 8
  },
});
console.log(result);

Python

import os
import requests

headers = {"Authorization": f"Key {os.environ['MRUN_API_KEY']}"}

submitted = requests.post(
    "https://queue.modelrunner.run/stability-ai/stable-audio-2.5/text-to-audio",
    headers=headers,
    json={
      "prompt": "gentle rain on a tin roof with distant rolling thunder and a soft wind, natural field-recording ambience",
      "seconds_total": 20,
      "guidance_scale": 1,
      "num_inference_steps": 8
    },
).json()

# Poll submitted["status_url"] until "COMPLETED", then:
result = requests.get(submitted["response_url"], headers=headers).json()

Input parameters

Input parameters of Stable Audio 2.5
NameTypeRequiredDescription
promptstringyesThe prompt to generate audio from. Describe genre, instrumentation, mood, and tempo for music, or the source, environment, and materials for sound effects.
seconds_totalintegernoDuration of the generated audio in seconds (1-190). Billing is a flat rate per generation regardless of length. Default: 190.
num_inference_stepsintegernoNumber of denoising steps. More steps can improve quality at the cost of speed. Default: 8.
guidance_scalenumbernoClassifier-free guidance scale; higher values follow the prompt more strictly. Default: 1.
seedintegernoRandom seed for reproducible generation. Leave empty for a random result.

Machine-readable: OpenAPI schema · llms.txt

Use Stable Audio 2.5 from Claude & Cursor (MCP)

Point Claude Code, Claude Desktop, Cursor, or any MCP client at the ModelRunner MCP server and Stable Audio 2.5 becomes a tool your assistant can call directly — it authorizes via OAuth (no API key in config) and runs this model with the run_model tool using the endpoint stability-ai/stable-audio-2.5/text-to-audio.

MCP client config (Claude Desktop, Cursor)

{
  "mcpServers": {
    "modelrunner": {
      "command": "npx",
      "args": ["-y", "mcp-remote", "https://mcp.modelrunner.run/mcp"]
    }
  }
}

Claude Code

claude mcp add --transport http modelrunner https://mcp.modelrunner.run/mcp

Then ask your assistant, for example: “Run stability-ai/stable-audio-2.5/text-to-audio on ModelRunner to generate music”. MCP setup guide.

Model Detayları

Model Detayları

Stable Audio 2.5, yazılı bir açıklamayı yaklaşık 190 saniyeye kadar süren, bitmiş bir ses klibine dönüştürür — bir müzik parçası, bir ses atmosferi, ortam sesi ya da bir ses efekti — ve sonucu WAV dosyası olarak döndürür. İstediğinizi `prompt` alanına yazın ("driving synthwave with a punchy kick and arpeggiated bass", "gentle rain on a window with distant thunder", "upbeat corporate acoustic bed, no vocals"), uzunluğu da `seconds_total` ile belirleyin. En güçlü yanı uzun soluklu üretimdir: birkaç saniyeyle sınırlı kısa ses efekti modellerinin aksine tek bir çağrı, tam bir fon müziği olarak kullanılabilecek dakikalarca süren parçalar üretebilir. Ayrıca ticari kullanım açısından güvenli çıktı için tamamen lisanslı bir veri setiyle eğitilmiştir.

## En uygun olduğu işler - Videolar, reklamlar, podcast’ler ve oyunlar için fon müziği ve enstrümantal altyapılar - Tek bir metin prompt’undan yaklaşık üç dakikaya kadar uzun soluklu parçalar ve loop’lar - Sahneler için ortam sesleri ve ses atmosferleri (yağmur, kafe uğultusu, orman, oda tonu) - Düz bir dille tarif edilen tek seferlik ses efektleri ve foley - Ticari açıdan güvenli, lisanslı veriyle eğitilmiş bir modelin önem taşıdığı, telif konusunda hassas ses işleri

## Şu durumlarda başka bir model seçin - Metinden üretmek yerine mevcut bir ses klibini dönüştürmek ya da yeniden tarzlamak istiyorsanız — bir audio-to-audio modeli kullanın - Doğal bir konuşma anlatımı ya da belirli bir seslendirme sesi gerekiyorsa — bir text-to-speech modeli kullanın - Yalnızca çok kısa, tek seferlik bir efekt gerekiyorsa ve minik kliplerde saniye başına faturalandırma istiyorsanız — saniye başına ücretlendirilen bir ses efekti modeli daha ucuza gelebilir

## İpuçları - `seconds_total` 1–190 saniye arasında değer alır; faturalandırma üretim başına sabit ücretle yapılır, yani uzun klipler kısalarla aynı fiyata gelir - Müzik için prompt’ta türü, enstrümantasyonu, ruh halini ve tempoyu; ses efektleri için sesin kaynağını, ortamı ve malzemeleri tarif edin - Temiz bir fon müziği istediğinizde prompt’a "no vocals" ya da "instrumental" yazın - Kaliteyi biraz artırmak için `num_inference_steps` değerini yükseltin (en fazla 8); prompt’a daha sıkı uyum için `guidance_scale` değerini yükseltin

## Sınırlamalar - Her çağrı tek bir WAV klibi döndürür; multitrack çıktı ya da stem ayırma yoktur - Çok kısa süreler, uzun kliplere göre müzikal açıdan daha az gelişmiş sonuçlar verebilir

ModelRunner JavaScript client ile çalıştırmak için: ```js import { modelrunner } from "@modelrunner/client";

const result = await modelrunner.subscribe("stability-ai/stable-audio-2.5/text-to-audio", { input: { prompt: "upbeat lofi hip hop instrumental with a warm vinyl texture", seconds_total: 60, }, }); ```