Catalog
Models
Featured models across image, video, audio and text.

BiRefNet Background Removal
zhengpeng7
Remove the background from an image and get a transparent-PNG cutout, with 11 selectable checkpoints and a soft alpha matte that keeps individual hair strands separated.

Kokoro-82M
hexgrad
Turn text into natural-sounding speech with 46 voices across 6 languages, billed by compute time at a fraction of a cent per clip.
Rodin Gen-2 by Hyper3D
hyper3d
Generate a clean, quad-topology 3D model with PBR textures (GLB, OBJ, USDZ, FBX, STL) from a text prompt or up to five reference images.

Wan 3.0 Image to Video
wan-video
Animate a still photo into a clip of up to 30 seconds at 480P, 720P or 1080P, optionally pinning a closing frame so the model generates the motion between two images, with a matching soundtrack included.

Qwen-Image 3.0
alibaba
Generate an image from a text prompt, with in-image text rendered natively in 12 languages and legible down to around 10px.

Wan 3.0 Text to Video
wan-video
Generate a video from a text prompt at 480P, 720P or 1080P, with a matching soundtrack included and clips running up to 30 seconds in one generation.

Happy Horse 1.1 Reference to Video
alibaba
Generate a 3-15 second video from up to nine reference photos, addressing each one positionally in the prompt so a specific character, product or prop stays recognizable in the shot.

Happy Horse 1.1 Image to Video
alibaba
Animate a still photo into a 3-15 second video at 720P or 1080P, with audio generated alongside the picture and the output frame shape taken straight from your image.

Gemini 3.1 Flash TTS
Turn text into expressive, directable speech in 30 voices — describe the delivery in plain language and get back a 24 kHz WAV.

GLM-5.2 Fast Preview
z-ai
Low-latency chat completions from the GLM-5.2 weights on throughput-tuned serving — consistently faster than the standard tier, at twice the price and identical answer quality.

DeepSeek V4 Pro
deepseek
Open-weight (MIT) thinking model at 1.6T parameters for hard reasoning and competitive-grade coding, with a 1,000,000-token context, tool calling and streaming over an OpenAI-compatible chat endpoint.

GLM-5.2
z-ai
Open-weight (MIT) thinking model for agentic coding and long-horizon reasoning, with a 1M-token context, seven levels of thinking effort, tool calling and streaming over an OpenAI-compatible chat endpoint.

Wan 2.7 Video Extend
wan-video
Extend an existing video clip into a longer one - up to 15 seconds in total, with matching audio, and optionally steered toward a closing frame you supply.

Wan 2.7 Image to Video (Audio Driven)
wan-video
Drive a still photo with your own audio clip: the track is used for lip-sync and action timing, producing a 2-15 second video at 720P or 1080P that performs in time with the sound.

Wan 2.7 Image to Video
wan-video
Animate a still photo into a 2-15 second video at 720P or 1080P, optionally pinning a closing frame, with background music or sound effects generated alongside the picture.

Gemini 3.5 Flash-Lite
The cheapest Gemini text tier — built for high-volume agentic tasks, translation and simple data processing over an OpenAI-compatible endpoint.

Gemini 3.7 Flash
Reasoning-first flagship text model for coding and agentic work, with tunable thinking levels, tool calling and streaming over an OpenAI-compatible endpoint.

Gemini 3.5 Flash
Fast, general-purpose text model served over an OpenAI-compatible chat completions endpoint, with tool calling, JSON mode and streaming.

Seedance 2.5 Video to Video
bytedance
Turn existing footage into a new video — restyle, restage, reframe or extend a clip you already have, steering it with up to 10 source videos named in the prompt.

Seedance 2.5 Reference to Video
bytedance
Generate a video steered by up to 30 reference images — composite a product, character, location, or style plate into one shot by naming each reference in the prompt.

Seedance 2.5 First & Last Frame
bytedance
Generate a video that starts on one image and ends on another — pin both ends of the shot and get a single take up to 30 seconds long with synchronized sound.

Seedance 2.5 Image to Video
bytedance
Animate a still photo into a video with synchronized audio — single takes up to 30 seconds long that keep your image's exact shape.

Seedance 2.5 Text to Video
bytedance
Generate a video with synchronized audio from a text prompt — single takes up to 30 seconds long, at 480p or 720p.

Happy Horse 1.1 Text to Video
alibaba
Generate a short video with synchronized native audio from a text prompt, with spoken dialogue lip-synced on screen, at 720P or 1080P.
Showing 1–24 of 167 models
