Ana içeriğe geç
minimax avatarı

MiniMax Speech-02 HD API

minimax/speech-02-hd

Metni 30+ dilde ve 300+ sesle doğal, yüksek sadakatli konuşmaya dönüştürün; duygu, hız, ses perdesi ve ses seviyesi kontrolü sizde.

üretilen ses için saniye başına 0.0015

Model girdisi

Input

The text to synthesize into speech. Up to 5000 characters.

The voice to speak with. 300+ voices supported; pass the voice id as a string. Examples: Wise_Woman, Friendly_Person, Deep_Voice_Man, Calm_Woman, Casual_Guy, Lively_Girl, Patient_Man.

Min: 0.5 - Max: 2

Speaking speed. 1 is normal; lower is slower, higher is faster.

Min: 0.01 - Max: 10

Output volume multiplier.

Min: -12 - Max: 12

Pitch shift in semitones. 0 is the voice's natural pitch.

Emotional tone of the delivery. Leave unset for a neutral read.

If true, normalize English text (e.g. numbers and units) before synthesis for more natural pronunciation.

Voice selection and delivery controls.

Additional Settings

Customize your input with more control.

Hint the primary language/dialect to improve recognition for non-default or mixed-language text. Use a language name or 'auto'. Leave unset to let the model detect it.

You need to be logged in to run this model and view results.
Log in

Model çıktısı

Output

Loading
Generated in 33.18 seconds
Logs (1 lines)

Örnek istekler

Örnekler

MiniMax Speech-02 HD API

MiniMax Speech-02 HD is a sound AI model by minimax. On ModelRunner it runs through a REST API or via MCP from any AI assistant, at $0.0015 per second of audio.

POST https://queue.modelrunner.run/minimax/speech-02-hd

cURL

# Submit a request to the queue. Input fields go at the top level of the
# body. The optional reserved "metadata" object holds your own flat string
# tags — stored on the request, never sent to the model; filter later with
# GET https://queue.modelrunner.run/requests?metadata=<url-encoded JSON>.
curl -X POST https://queue.modelrunner.run/minimax/speech-02-hd \
  -H "Authorization: Key $MRUN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "Welcome to ModelRunner. This is a live verification of the MiniMax Speech-02 HD text to speech model.",
    "metadata": {
      "project": "my-project"
    }
  }'
# → { "request_id": "...", "status_url": "...", "response_url": "..." }

# Poll status_url until "COMPLETED", then fetch the result
curl "https://queue.modelrunner.run/minimax/speech-02-hd/requests/$REQUEST_ID/status" \
  -H "Authorization: Key $MRUN_API_KEY"
curl "https://queue.modelrunner.run/minimax/speech-02-hd/requests/$REQUEST_ID" \
  -H "Authorization: Key $MRUN_API_KEY"

JavaScript

import { modelrunner } from "@modelrunner/client";

const result = await modelrunner.subscribe("minimax/speech-02-hd", {
  input: {
    "text": "Welcome to ModelRunner. This is a live verification of the MiniMax Speech-02 HD text to speech model."
  },
});
console.log(result);

Python

import os
import requests

headers = {"Authorization": f"Key {os.environ['MRUN_API_KEY']}"}

submitted = requests.post(
    "https://queue.modelrunner.run/minimax/speech-02-hd",
    headers=headers,
    json={
      "text": "Welcome to ModelRunner. This is a live verification of the MiniMax Speech-02 HD text to speech model."
    },
).json()

# Poll submitted["status_url"] until "COMPLETED", then:
result = requests.get(submitted["response_url"], headers=headers).json()

Input parameters

Input parameters of MiniMax Speech-02 HD
NameTypeRequiredDescription
textstringyesThe text to synthesize into speech. Up to 5000 characters.
voice_settingobjectnoVoice selection and delivery controls.
language_boostenumnoHint the primary language/dialect to improve recognition for non-default or mixed-language text. Use a language name or 'auto'. Leave unset to let the model detect it. One of: Chinese, Chinese,Yue, English, Arabic, Russian, Spanish, French, Portuguese, German, Turkish, Dutch, Ukrainian, Vietnamese, Indonesian, Japanese, Italian, Korean, Thai, Polish, Romanian, Greek, Czech, Finnish, Hindi, Bulgarian, Danish, Hebrew, Malay, Slovak, Swedish, Croatian, Hungarian, Norwegian, Slovenian, Catalan, Nynorsk, Afrikaans, auto.

Machine-readable: OpenAPI schema · llms.txt

Use MiniMax Speech-02 HD from Claude & Cursor (MCP)

Point Claude Code, Claude Desktop, Cursor, or any MCP client at the ModelRunner MCP server and MiniMax Speech-02 HD becomes a tool your assistant can call directly — it authorizes via OAuth (no API key in config) and runs this model with the run_model tool using the endpoint minimax/speech-02-hd.

MCP client config (Claude Desktop, Cursor)

{
  "mcpServers": {
    "modelrunner": {
      "command": "npx",
      "args": ["-y", "mcp-remote", "https://mcp.modelrunner.run/mcp"]
    }
  }
}

Claude Code

claude mcp add --transport http modelrunner https://mcp.modelrunner.run/mcp

Then ask your assistant, for example: “Run minimax/speech-02-hd on ModelRunner to generate sound”. MCP setup guide.

Model Detayları

Model Detayları

MiniMax Speech-02 HD, yazılı metni doğal, stüdyo kalitesinde konuşmaya dönüştürür ve barındırılan bir MP3 döndürür. Seslendirilmesini istediğiniz metni verip bir ses seçersiniz; tonlamayı, tempoyu ve telaffuzu model üstlenir, duygu, hız, ses perdesi ve ses seviyesi üzerinde ince kontrol sunar. 30+ dili ve 300+ hazır sesi desteklemesi onu anlatım, seslendirme, IVR/telefon anonsları, sesli kitaplar ve karakter diyalogları için güçlü bir varsayılan tercih haline getirir. HD profili netliği ve sadakati öncelediğinden, insanların baştan sona dinleyeceği içeriklere çok uygundur.

## En uygun olduğu işler - Videolar, açıklayıcı içerikler ve reklamlar için anlatım ve seslendirme - Tutarlı, net bir okuyuşun önemli olduğu sesli kitaplar ve uzun metinler - IVR, telefon menüleri ve otomatik anonslar - Duygu kontrolüyle karakter diyalogları ve canlandırmalı okumalar - Çok dilli içerik — 30+ dilde aynı pipeline

## Şu durumlarda başka bir model seçin - Melodisi ve enstrümantasyonu olan bir müzik parçası gerekiyorsa — bir müzik üretim modeli kullanın - Konuşma yerine genel ses efektleri ya da ortam sesi istiyorsanız — bir text-to-audio ses efekti modeli kullanın - Canlı bir çağrıda gerçek zamanlı, ultra düşük latency gerektiren streaming TTS gerekiyorsa — bu model token stream değil, bitmiş bir dosya döndürür

## İpuçları - Konuşmacıyı seçmek için `voice_setting.voice_id` alanını ayarlayın. Varsayılan `Wise_Woman`; diğer örnekler arasında `Friendly_Person`, `Deep_Voice_Man`, `Calm_Woman`, `Casual_Guy`, `Lively_Girl` ve `Patient_Man` var (300+ ses desteklenir — ses id’sini string olarak verin). - Okuyuşu renklendirmek için `voice_setting.emotion` kullanın (`happy`, `sad`, `angry`, `fearful`, `disgusted`, `surprised`, `neutral`); nötr bir okuyuş için boş bırakın. - Tempoyu ve tonu `voice_setting.speed` (0.5–2.0), `voice_setting.pitch` (-12 ile 12 arası) ve `voice_setting.vol` (0.01–10) ile ayarlayın. - Girdi metnini okunmasını istediğiniz gibi noktalayın — virgüller ve noktalar duraklamaları ve tonlamayı belirler.

## Gelişmiş Yapılandırma - `voice_setting.english_normalization` (varsayılan `false`): `true` olduğunda, daha doğal bir telaffuz için sentezden önce İngilizce metni (ör. sayılar ve birimler) normalleştirir; karşılığında küçük bir latency maliyeti olur. API üzerinden ayarlanır. - `language_boost` (varsayılan olarak boş): varsayılan dışı ya da karışık dilli metinlerde tanımayı iyileştirmek için birincil dili/lehçeyi belirtin. Bir dil adı (ör. `English`, `Spanish`, `Japanese`, `Arabic`) ya da `auto` kabul eder.

ModelRunner JavaScript client ile çalıştırmak için: ```js import { modelrunner } from "@modelrunner/client";

const result = await modelrunner.subscribe("minimax/speech-02-hd", { input: { text: "Welcome to ModelRunner. This is high-definition text to speech.", voice_setting: { voice_id: "Wise_Woman", speed: 1, emotion: "happy" }, }, }); ```