Skip to main content

Catalog

Models

Featured models across image, video, audio and text.

BiRefNet Background Removal

BiRefNet Background Removal

zhengpeng7

Remove the background from an image and get a transparent-PNG cutout, with 11 selectable checkpoints and a soft alpha matte that keeps individual hair strands separated.

imageimageremove backgroundmask
Kokoro-82M

Kokoro-82M

hexgrad

Turn text into natural-sounding speech with 46 voices across 6 languages, billed by compute time at a fraction of a cent per clip.

textaudio
Rodin Gen-2 by Hyper3D

Rodin Gen-2 by Hyper3D

hyper3d

Generate a clean, quad-topology 3D model with PBR textures (GLB, OBJ, USDZ, FBX, STL) from a text prompt or up to five reference images.

image-to-3d
Wan 3.0 Image to Video

Wan 3.0 Image to Video

wan-video

Animate a still photo into a clip of up to 30 seconds at 480P, 720P or 1080P, optionally pinning a closing frame so the model generates the motion between two images, with a matching soundtrack included.

imagevideo
Qwen-Image 3.0

Qwen-Image 3.0

alibaba

Generate an image from a text prompt, with in-image text rendered natively in 12 languages and legible down to around 10px.

textimage
Wan 3.0 Text to Video

Wan 3.0 Text to Video

wan-video

Generate a video from a text prompt at 480P, 720P or 1080P, with a matching soundtrack included and clips running up to 30 seconds in one generation.

textvideo
Happy Horse 1.1 Reference to Video

Happy Horse 1.1 Reference to Video

alibaba

Generate a 3-15 second video from up to nine reference photos, addressing each one positionally in the prompt so a specific character, product or prop stays recognizable in the shot.

imagevideo
Happy Horse 1.1 Image to Video

Happy Horse 1.1 Image to Video

alibaba

Animate a still photo into a 3-15 second video at 720P or 1080P, with audio generated alongside the picture and the output frame shape taken straight from your image.

imagevideo
Gemini 3.1 Flash TTS

Gemini 3.1 Flash TTS

google

Turn text into expressive, directable speech in 30 voices — describe the delivery in plain language and get back a 24 kHz WAV.

textaudio
GLM-5.2 Fast Preview

GLM-5.2 Fast Preview

z-ai

Low-latency chat completions from the GLM-5.2 weights on throughput-tuned serving — consistently faster than the standard tier, at twice the price and identical answer quality.

texttext
DeepSeek V4 Pro

DeepSeek V4 Pro

deepseek

Open-weight (MIT) thinking model at 1.6T parameters for hard reasoning and competitive-grade coding, with a 1,000,000-token context, tool calling and streaming over an OpenAI-compatible chat endpoint.

texttext
GLM-5.2

GLM-5.2

z-ai

Open-weight (MIT) thinking model for agentic coding and long-horizon reasoning, with a 1M-token context, seven levels of thinking effort, tool calling and streaming over an OpenAI-compatible chat endpoint.

texttext
Wan 2.7 Video Extend

Wan 2.7 Video Extend

wan-video

Extend an existing video clip into a longer one - up to 15 seconds in total, with matching audio, and optionally steered toward a closing frame you supply.

videovideoextend
Wan 2.7 Image to Video (Audio Driven)

Wan 2.7 Image to Video (Audio Driven)

wan-video

Drive a still photo with your own audio clip: the track is used for lip-sync and action timing, producing a 2-15 second video at 720P or 1080P that performs in time with the sound.

imagevideo
Wan 2.7 Image to Video

Wan 2.7 Image to Video

wan-video

Animate a still photo into a 2-15 second video at 720P or 1080P, optionally pinning a closing frame, with background music or sound effects generated alongside the picture.

imagevideo
Gemini 3.5 Flash-Lite

Gemini 3.5 Flash-Lite

google

The cheapest Gemini text tier — built for high-volume agentic tasks, translation and simple data processing over an OpenAI-compatible endpoint.

texttext
Gemini 3.7 Flash

Gemini 3.7 Flash

google

Reasoning-first flagship text model for coding and agentic work, with tunable thinking levels, tool calling and streaming over an OpenAI-compatible endpoint.

texttext
Gemini 3.5 Flash

Gemini 3.5 Flash

google

Fast, general-purpose text model served over an OpenAI-compatible chat completions endpoint, with tool calling, JSON mode and streaming.

texttext
Seedance 2.5 Video to Video

Seedance 2.5 Video to Video

bytedance

Turn existing footage into a new video — restyle, restage, reframe or extend a clip you already have, steering it with up to 10 source videos named in the prompt.

videovideoeditextend
Seedance 2.5 Reference to Video

Seedance 2.5 Reference to Video

bytedance

Generate a video steered by up to 30 reference images — composite a product, character, location, or style plate into one shot by naming each reference in the prompt.

imagevideo
Seedance 2.5 First & Last Frame

Seedance 2.5 First & Last Frame

bytedance

Generate a video that starts on one image and ends on another — pin both ends of the shot and get a single take up to 30 seconds long with synchronized sound.

imagevideo
Seedance 2.5 Image to Video

Seedance 2.5 Image to Video

bytedance

Animate a still photo into a video with synchronized audio — single takes up to 30 seconds long that keep your image's exact shape.

imagevideo
Seedance 2.5 Text to Video

Seedance 2.5 Text to Video

bytedance

Generate a video with synchronized audio from a text prompt — single takes up to 30 seconds long, at 480p or 720p.

textvideo
Happy Horse 1.1 Text to Video

Happy Horse 1.1 Text to Video

alibaba

Generate a short video with synchronized native audio from a text prompt, with spoken dialogue lip-synced on screen, at 720P or 1080P.

textvideo

Showing 124 of 167 models