Audio & Music Models
Generate music, speech, sound effects, and transcriptions with AI. Compare top audio models for text-to-music, text-to-speech, and speech-to-text.

ACE-Step
ace-studio
Generate full songs or instrumental music from genre tags and optional lyrics, with duration you control up to 4 minutes.
musicmeta / musicgen
A fast, controllable auto-regressive Transformer for high-fidelity music generation.
soundMiniMax Speech-02 HD
minimax
Turn text into natural, high-fidelity speech in 30+ languages with 300+ voices plus emotion, speed, pitch, and volume control.
sound
LTX-2.3 Text-to-Audio
lightricks
Generate sound effects, ambience, and spoken-style audio from a text prompt, with duration you control down to the frame.
soundElevenLabs Scribe v1
elevenlabs
Transcribe speech audio into accurate text with word-level timestamps, speaker labels, and audio-event tags across 99 languages.
speech-to-textLyria 2
Generate ~30 seconds of high-fidelity instrumental music from a text prompt, as a 48kHz WAV file.
musicElevenLabs Sound Effects V2
elevenlabs
Generate sound effects, Foley, and ambience from a text prompt, returning a hosted MP3.
sound
Whisper Large v3
openai
Transcribe or translate speech audio into text across 99 languages, with segment/word timestamps and optional speaker diarization.
speech-to-textChatterbox TTS
resemble-ai
Turn text into expressive speech and clone any voice from a short reference recording, with fine control over emotional intensity.
soundElevenLabs Multilingual v2
elevenlabs
Turn text into natural, expressive speech in 29 languages with ElevenLabs Multilingual v2 voices, with controls for stability, similarity, and style.
soundDeepFilterNet 3
rikorose
Clean up a noisy speech recording by removing background noise and upsampling it to studio-quality 48 kHz audio.
audio-to-audioSAM Audio — Separate
meta
Isolate any sound from an audio mixture by describing it in plain language.
audio-to-audioElevenLabs Audio Isolation
elevenlabs
Strip background noise and music from a recording to isolate clean, studio-quality speech, returning the isolated voice as an MP3.
audio-to-audio