Skip to main content

Catalog

Models

Featured models across image, video, audio and text.

Lyria 3 Clip

Lyria 3 Clip

google

Generate a 30-second song clip — vocals and lyrics included by default — from a single text prompt, as an MP3.

textaudio
Stable Audio 2.5 Audio-to-Audio

Stable Audio 2.5 Audio-to-Audio

stability-ai

Transform an existing audio clip into new music or sound effects guided by a text prompt — restyle, re-instrument, or reimagine a source track, returned as a WAV.

audioaudio
Chatterbox HD TTS

Chatterbox HD TTS

resemble-ai

Turn text into high-definition speech with nine named voices or a cloned voice, plus an optional 48kHz upscale toggle for higher-fidelity audio.

textaudio
Chatterbox Multilingual TTS

Chatterbox Multilingual TTS

resemble-ai

Turn text into natural speech across 23 languages with a language selector and emotion-intensity control.

textaudio
Stable Audio 2.5

Stable Audio 2.5

stability-ai

Generate long-form music and sound effects from a text prompt — up to ~190 seconds of WAV audio in a single call.

textaudio
SAM Audio — Separate

SAM Audio — Separate

meta

Isolate any sound from an audio mixture by describing it in plain language.

audioaudiosegmentation
Lyria 2

Lyria 2

google

Generate ~30 seconds of high-fidelity instrumental music from a text prompt, as a 48kHz WAV file.

textaudio
ElevenLabs Sound Effects V2

ElevenLabs Sound Effects V2

elevenlabs

Generate sound effects, Foley, and ambience from a text prompt, returning a hosted MP3.

textaudio
Chatterbox TTS

Chatterbox TTS

resemble-ai

Turn text into expressive speech and clone any voice from a short reference recording, with fine control over emotional intensity.

textaudio
ElevenLabs Multilingual v2

ElevenLabs Multilingual v2

elevenlabs

Turn text into natural, expressive speech in 29 languages with ElevenLabs Multilingual v2 voices, with controls for stability, similarity, and style.

textaudio
ACE-Step

ACE-Step

ace-studio

Generate full songs or instrumental music from genre tags and optional lyrics, with duration you control up to 4 minutes.

textaudio
MiniMax Speech-02 HD

MiniMax Speech-02 HD

minimax

Turn text into natural, high-fidelity speech in 30+ languages with 300+ voices plus emotion, speed, pitch, and volume control.

textaudio
LTX-2.3 Text-to-Audio

LTX-2.3 Text-to-Audio

lightricks

Generate sound effects, ambience, and spoken-style audio from a text prompt, with duration you control down to the frame.

textaudio
musicgen

musicgen

meta

A fast, controllable auto-regressive Transformer for high-fidelity music generation.

textaudio

Showing 114 of 14 models