Skip to main content
Back to Explore

Audio & Music Models

Generate music, speech, sound effects, and transcriptions with AI. Compare top audio models for text-to-music, text-to-speech, and speech-to-text.

ACE-Step

ACE-Step

ace-studio

Generate full songs or instrumental music from genre tags and optional lyrics, with duration you control up to 4 minutes.

music
musicgen

meta / musicgen

A fast, controllable auto-regressive Transformer for high-fidelity music generation.

sound
MiniMax Speech-02 HD

MiniMax Speech-02 HD

minimax

Turn text into natural, high-fidelity speech in 30+ languages with 300+ voices plus emotion, speed, pitch, and volume control.

sound
LTX-2.3 Text-to-Audio

LTX-2.3 Text-to-Audio

lightricks

Generate sound effects, ambience, and spoken-style audio from a text prompt, with duration you control down to the frame.

sound
ElevenLabs Scribe v1

ElevenLabs Scribe v1

elevenlabs

Transcribe speech audio into accurate text with word-level timestamps, speaker labels, and audio-event tags across 99 languages.

speech-to-text
Lyria 2

Lyria 2

google

Generate ~30 seconds of high-fidelity instrumental music from a text prompt, as a 48kHz WAV file.

music
ElevenLabs Sound Effects V2

ElevenLabs Sound Effects V2

elevenlabs

Generate sound effects, Foley, and ambience from a text prompt, returning a hosted MP3.

sound
Whisper Large v3

Whisper Large v3

openai

Transcribe or translate speech audio into text across 99 languages, with segment/word timestamps and optional speaker diarization.

speech-to-text
Chatterbox TTS

Chatterbox TTS

resemble-ai

Turn text into expressive speech and clone any voice from a short reference recording, with fine control over emotional intensity.

sound
ElevenLabs Multilingual v2

ElevenLabs Multilingual v2

elevenlabs

Turn text into natural, expressive speech in 29 languages with ElevenLabs Multilingual v2 voices, with controls for stability, similarity, and style.

sound
DeepFilterNet 3

DeepFilterNet 3

rikorose

Clean up a noisy speech recording by removing background noise and upsampling it to studio-quality 48 kHz audio.

audio-to-audio
SAM Audio — Separate

SAM Audio — Separate

meta

Isolate any sound from an audio mixture by describing it in plain language.

audio-to-audio
ElevenLabs Audio Isolation

ElevenLabs Audio Isolation

elevenlabs

Strip background noise and music from a recording to isolate clean, studio-quality speech, returning the isolated voice as an MP3.

audio-to-audio