Unlimited Open Source Models

Get Plan
Skip to main content

AI APIs for Developers

Filter by use case

AI Model APIs

205 models
LTX 2.5 Pro Text To Video

ltx

LTX 2.5 Pro Text To Video

LTX 2.5 Pro turns text prompts into cinematic video with native synchronized audio — dialogue, ambience, and sound effects generated alongside the frames. Output at 720p or 1080p, 25 or 50 fps, up to 10 seconds, from a single ModelsLab API call.

Closed SourceFilmmaker GradeCinematic+2
LTX 2.5 Pro Image To Video

ltx

LTX 2.5 Pro Image To Video

LTX 2.5 Pro Image to Video animates a single still into a cinematic clip, using a text prompt to direct motion, camera movement, and atmosphere. Native synchronized audio, 720p or 1080p output, 25 or 50 fps, and 6 to 10 second durations — all from one end

Closed SourceFilmmaker GradeCinematic+2
Grok Imagine Image 2.0 Image Edit

xAI

Grok Imagine Image 2.0 Image Edit

Grok Imagine Image 2.0 Image Editing modifies existing images from plain text instructions — add, remove, restyle, or merge subjects while preserving the original look. Handles up to 14 input images in a single generation across seven aspect ratios.

Closed SourceNew Added2K Output+2
Grok Imagine Image 2.0 Text To Image

xAI

Grok Imagine Image 2.0 Text To Image

Grok Imagine Image 2.0 turns text prompts into 2K images with stronger prompt adherence, cleaner in-image text, and sharper detail than the original Grok Imagine. Five aspect ratios, 1K or 2K output, one ModelsLab endpoint.

Closed SourceNew Added2K Output+2
Seedance 2.5 Text to Video

Bytedance

Seedance 2.5 Text to Video

Seedance 2.5 writes a full scene from one text prompt — up to 30 seconds of continuous video with synced dialogue, ambience and score, no stitching required. Prompt in 11 languages, pick any aspect ratio from 21:9 to 9:16, and render at 480p or 720p

Closed SourceNew AddedNative Sync Audio+2
Seedance 2.5 Image To Video

Bytedance

Seedance 2.5 Image To Video

Seedance 2.5 animates a single still into a video up to 30 seconds long, with native sound and dialogue in 11 languages. Set a first frame, or pin both first and last frames to control exactly where the shot starts and ends. Outputs 480p or 720p.

Closed SourceNew AddedNative Sync Audio+2
Seedance 2.5 Multimodal Reference to Video

Bytedance

Seedance 2.5 Multimodal Reference to Video

Seedance 2.5 turns up to 50 multimodal references 30 images, 10 video clips and 10 audio tracks into one coherent video up to 30 seconds long, with native audio in 11 languages. Lock characters, products, motion and sound in a single call at 480p or 720p.

Closed SourceNew AddedNative Sync Audio+2
Qwen Image 3.0 Pro Image Edit

Alibaba

Qwen Image 3.0 Pro Image Edit

Qwen Image 3.0 Pro edits images from plain-text instructions and up to 3 reference images. Swap outfits, change scenes, restyle a shot or blend subjects while keeping faces and details intact.

Closed SourceNew AddedBest Selling+2
Qwen Image 3.0 Pro Text To Image

Alibaba

Qwen Image 3.0 Pro Text To Image

Qwen Image 3.0 Pro turns a single text prompt into a finished PNG at up to 2048×2048. Built-in prompt rewriting sharpens short prompts, negative prompts strip what you don't want, and seeds keep results repeatable.

Closed SourceNew AddedBest Selling+2
Flux 3 Video To Video

Black Forest Labs

Flux 3 Video To Video

FLUX 3 in video-continuation mode. Upload an MP4 and FLUX 3 carries the shot on from its final frames for another 5-20 seconds, audio included.

Closed SourceBest for CreatorsNative Sync Audio+2
Flux 3 Image To Video

Black Forest Labs

Flux 3 Image To Video

FLUX 3 in image-to-video mode. Drop in 1-10 images as keyframes, set the timing, and get a 5-20s HD clip with synchronized audio.

Closed SourceBest for CreatorsNative Sync Audio+2
Flux 3 Text To Video

Black Forest Labs

Flux 3 Text To Video

Black Forest Labs' FLUX 3 in text-to-video mode. Type a prompt, get a 5-20s HD or FHD clip with synchronized audio built in

Closed SourceBest for CreatorsNative Sync Audio+2
MiniMax H3 Start/ End Frame To Video

Minmax

MiniMax H3 Start/ End Frame To Video

Give H3 a start frame and an end frame — it generates the transition in native 2K with sound. Deterministic in and out points for edit-ready clips.

Closed Source2K Output15 sec Output+3
MiniMax H3 Image To Video

Minmax

MiniMax H3 Image To Video

Feed one image and a prompt, get a native 2K clip with sound. Preserves subject identity, lighting, and on-image text across motion.

Closed Source2K Output15 sec Output+3
MiniMax-H3 Reference To Video

Minmax

MiniMax-H3 Reference To Video

Pass reference images, video, or audio and describe how they relate. H3 handles subject, style, and motion transfer in one unified call.

Closed Source2K Output15 sec Output+3
MiniMax H3 Text to Video (Hailuo-03)

Minmax

MiniMax H3 Text to Video (Hailuo-03)

MiniMax H3 is a frontier AI video model that turns text prompts into stunning 2K videos in just seconds. Generate 5–15 second cinematic clips in seven aspect ratios, perfect for social media, marketing, and professional content.

Closed Source2K Output15 sec Output+3