Ana içeriğe geç
stability-ai avatarı

Stable Audio 2.5 Audio-to-Audio API

stability-ai/stable-audio-2.5/audio-to-audio

Mevcut bir ses klibini metin prompt’uyla yeni bir müziğe ya da ses efektine dönüştürün — kaynak parçayı yeniden tarzlayın, enstrümantasyonunu değiştirin ya da yeniden yorumlayın; sonuç WAV olarak döner.

0.2

Model girdisi

Input

The text prompt guiding the transformation. Describe the genre, instrumentation, mood, and tempo you want the source audio reshaped into.

The source audio clip to transform (accepted formats: mp3, ogg, wav, m4a, aac).

Additional Settings

Customize your input with more control.

Min: 0.01 - Max: 1

How much the source audio is transformed: near 0 keeps it almost identical to the input, near 1 ignores the input and follows only the prompt.

Min: 4 - Max: 8

Number of denoising steps. More steps can improve quality at the cost of speed.

Min: 1 - Max: 190

Duration of the generated audio in seconds (1-190). Defaults to the source clip's length if unset. Billing is a flat rate per generation regardless of length.

Min: 1 - Max: 25

Classifier-free guidance scale; higher values follow the prompt more strictly.

Random seed for reproducible generation. Leave empty for a random result.

You need to be logged in to run this model and view results.
Log in

Model çıktısı

Output

Loading
Generated in 6.337 seconds
Logs (1 lines)

Örnek istekler

Örnekler

Stable Audio 2.5 Audio-to-Audio API

Stable Audio 2.5 Audio-to-Audio is a audio-to-audio AI model by stability-ai. On ModelRunner it runs through a REST API or via MCP from any AI assistant, at $0.2 per audio clip.

POST https://queue.modelrunner.run/stability-ai/stable-audio-2.5/audio-to-audio

cURL

# Submit a request to the queue. Input fields go at the top level of the
# body. The optional reserved "metadata" object holds your own flat string
# tags — stored on the request, never sent to the model; filter later with
# GET https://queue.modelrunner.run/requests?metadata=<url-encoded JSON>.
curl -X POST https://queue.modelrunner.run/stability-ai/stable-audio-2.5/audio-to-audio \
  -H "Authorization: Key $MRUN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "reimagine as a mellow lofi hip-hop beat with warm rhodes piano, vinyl crackle, and a laid-back swing",
    "strength": 0.8,
    "audio_url": "https://media.modelrunner.ai/B5FtekTs4xvdeT653Nl7R.wav",
    "total_seconds": 30,
    "guidance_scale": 1,
    "num_inference_steps": 8,
    "metadata": {
      "project": "my-project"
    }
  }'
# → { "request_id": "...", "status_url": "...", "response_url": "..." }

# Poll status_url until "COMPLETED", then fetch the result
curl "https://queue.modelrunner.run/stability-ai/stable-audio-2.5/audio-to-audio/requests/$REQUEST_ID/status" \
  -H "Authorization: Key $MRUN_API_KEY"
curl "https://queue.modelrunner.run/stability-ai/stable-audio-2.5/audio-to-audio/requests/$REQUEST_ID" \
  -H "Authorization: Key $MRUN_API_KEY"

JavaScript

import { modelrunner } from "@modelrunner/client";

const result = await modelrunner.subscribe("stability-ai/stable-audio-2.5/audio-to-audio", {
  input: {
    "prompt": "reimagine as a mellow lofi hip-hop beat with warm rhodes piano, vinyl crackle, and a laid-back swing",
    "strength": 0.8,
    "audio_url": "https://media.modelrunner.ai/B5FtekTs4xvdeT653Nl7R.wav",
    "total_seconds": 30,
    "guidance_scale": 1,
    "num_inference_steps": 8
  },
});
console.log(result);

Python

import os
import requests

headers = {"Authorization": f"Key {os.environ['MRUN_API_KEY']}"}

submitted = requests.post(
    "https://queue.modelrunner.run/stability-ai/stable-audio-2.5/audio-to-audio",
    headers=headers,
    json={
      "prompt": "reimagine as a mellow lofi hip-hop beat with warm rhodes piano, vinyl crackle, and a laid-back swing",
      "strength": 0.8,
      "audio_url": "https://media.modelrunner.ai/B5FtekTs4xvdeT653Nl7R.wav",
      "total_seconds": 30,
      "guidance_scale": 1,
      "num_inference_steps": 8
    },
).json()

# Poll submitted["status_url"] until "COMPLETED", then:
result = requests.get(submitted["response_url"], headers=headers).json()

Input parameters

Input parameters of Stable Audio 2.5 Audio-to-Audio
NameTypeRequiredDescription
promptstringyesThe text prompt guiding the transformation. Describe the genre, instrumentation, mood, and tempo you want the source audio reshaped into.
audio_urlstring (uri)yesThe source audio clip to transform (accepted formats: mp3, ogg, wav, m4a, aac).
strengthnumbernoHow much the source audio is transformed: near 0 keeps it almost identical to the input, near 1 ignores the input and follows only the prompt. Default: 0.8.
num_inference_stepsintegernoNumber of denoising steps. More steps can improve quality at the cost of speed. Default: 8.
total_secondsintegernoDuration of the generated audio in seconds (1-190). Defaults to the source clip's length if unset. Billing is a flat rate per generation regardless of length.
guidance_scalenumbernoClassifier-free guidance scale; higher values follow the prompt more strictly. Default: 1.
seedintegernoRandom seed for reproducible generation. Leave empty for a random result.

Machine-readable: OpenAPI schema · llms.txt

Use Stable Audio 2.5 Audio-to-Audio from Claude & Cursor (MCP)

Point Claude Code, Claude Desktop, Cursor, or any MCP client at the ModelRunner MCP server and Stable Audio 2.5 Audio-to-Audio becomes a tool your assistant can call directly — it authorizes via OAuth (no API key in config) and runs this model with the run_model tool using the endpoint stability-ai/stable-audio-2.5/audio-to-audio.

MCP client config (Claude Desktop, Cursor)

{
  "mcpServers": {
    "modelrunner": {
      "command": "npx",
      "args": ["-y", "mcp-remote", "https://mcp.modelrunner.run/mcp"]
    }
  }
}

Claude Code

claude mcp add --transport http modelrunner https://mcp.modelrunner.run/mcp

Then ask your assistant, for example: “Run stability-ai/stable-audio-2.5/audio-to-audio on ModelRunner to generate audio”. MCP setup guide.

Model Detayları

Model Detayları

Stable Audio 2.5 Audio-to-Audio, verdiğiniz bir kaynak ses klibini metin açıklamasının yönlendirmesiyle yeni bir klibe dönüştürür. `audio_url` ile bir parça yükleyin, istediğiniz sesi `prompt` alanında tarif edin ("bunu zengin, sinematik bir orkestra aranjmanına dönüştür", "enerjik bir elektronik dans parçasına remix’le", "plak cızırtılı bir lofi hip-hop beat’i olarak yeniden yorumla"); model size yepyeni bir WAV klibi döndürür. Orijinalden ne kadar uzaklaşacağını `strength` ile ayarlayın — düşük değerler kaynağa yakın kalır, yüksek değerler prompt’u daha serbestçe izler. Text-to-audio sürümüyle aynı, lisanslı veriyle eğitilmiş ve ticari kullanım açısından güvenli Stable Audio 2.5 modeli üzerine kuruludur: hem müzik hem de ses efektleriyle çalışır ve yaklaşık 190 saniyeye kadar klip üretebilir.

## En uygun olduğu işler - Mevcut bir parçayı yeniden tarzlamak ya da enstrümantasyonunu değiştirmek (folk gitar → orkestra, akustik → elektronik) - Kaba bir müzikal fikri ya da mırıldanılmış bir melodiyi daha dolgun bir aranjmana dönüştürmek - Kaynağın yapısını koruyarak ortam seslerini ya da ses efektlerini yeni bir tarzda yeniden yorumlamak - A/B seçenekleri için bir referans klibin prompt’la yönlendirilen varyasyonlarını üretmek - Lisanslı veriyle eğitilmiş bir modelin önem taşıdığı, ticari açıdan güvenli ses işleri

## Şu durumlarda başka bir model seçin - Kaynak klip olmadan metinden ses üretmek istiyorsanız — Stable Audio 2.5 text-to-audio modelini kullanın - Sözleri koruyarak konuşan kişinin sesini değiştirmek istiyorsanız — bir speech-to-speech / ses dönüştürme modeli kullanın - Belirli bir seste temiz bir anlatım gerekiyorsa — bir text-to-speech modeli kullanın - Yalnızca mevcut bir kaydın gürültüsünü gidermeniz ya da içinden bir sesi ayırmanız gerekiyorsa — bir ses temizleme / izolasyon modeli kullanın

## İpuçları - `strength` (0.01–1, varsayılan 0.8) en önemli ayardır: ~0.3–0.5 kaynağı tanınır halde bırakır, ~0.8–1.0 prompt’a ağırlık verir - `total_seconds` (1–190) çıktı uzunluğunu belirler; kaynak klibin süresine uymak için boş bırakın - Müzikal dönüşümlerde prompt’ta türü, enstrümantasyonu, ruh halini ve tempoyu tarif edin - Kaliteyi biraz artırmak için `num_inference_steps` değerini yükseltin (en fazla 8); prompt’a daha sıkı uyum için `guidance_scale` değerini yükseltin - Kabul edilen kaynak formatları: mp3, ogg, wav, m4a, aac

## Sınırlamalar - Her çağrı tek bir WAV klibi döndürür; multitrack çıktı ya da stem ayırma yoktur - Çok yüksek `strength` değeri kaynağı fiilen yok sayar ve text-to-audio üretimi gibi davranır

ModelRunner JavaScript client ile çalıştırmak için: ```js import { modelrunner } from "@modelrunner/client";

const result = await modelrunner.subscribe("stability-ai/stable-audio-2.5/audio-to-audio", { input: { prompt: "transform into a lush cinematic orchestral arrangement with sweeping strings", audio_url: "https://example.com/source-guitar.wav", strength: 0.7, }, }); ```