Catalog
Models
Featured models across image, video, audio and text.
Stable Audio 2.5 Audio-to-Audio
stability-ai
Transform an existing audio clip into new music or sound effects guided by a text prompt — restyle, re-instrument, or reimagine a source track, returned as a WAV.
Chatterbox HD Voice Conversion
resemble-ai
Convert a spoken recording into a different target voice while keeping the original words and delivery, with nine named HD voices and an optional 48kHz upscale for higher-fidelity output.
Chatterbox HD TTS
resemble-ai
Turn text into high-definition speech with nine named voices or a cloned voice, plus an optional 48kHz upscale toggle for higher-fidelity audio.
Chatterbox Voice Conversion
resemble-ai
Convert a spoken recording into a different target voice while keeping the original words, timing, and delivery — a speech-to-speech voice changer.
DeepFilterNet 3
rikorose
Clean up a noisy speech recording by removing background noise and upsampling it to studio-quality 48 kHz audio.
SAM Audio — Separate
meta
Isolate any sound from an audio mixture by describing it in plain language.
Chatterbox TTS
resemble-ai
Turn text into expressive speech and clone any voice from a short reference recording, with fine control over emotional intensity.
ElevenLabs Audio Isolation
elevenlabs
Strip background noise and music from a recording to isolate clean, studio-quality speech, returning the isolated voice as an MP3.
musicgen
meta
A fast, controllable auto-regressive Transformer for high-fidelity music generation.
Showing 1–9 of 9 models
