Access the best AI models: FLUX, Nano Banana, Kling, Seedance, GPT-5, Claude 5, ElevenLabs & more. Free to try!
Start with your task
From your first idea to something worth sharing.
Starting prices shown. Review the exact cost before you run.
240 models found

MMAudio Foley
200 +
Adds AI-generated sound effects, foley and ambience to a silent video, synchronized to the visual content. Describe the desired sound in the prompt.
Supports references
Generates audio

Qwen3.8 Omni Flash
10 +
Qwen3.8 Omni Flash is Alibaba's first Qwen model built around agentic capabilities with native omni-modal audio-video understanding, 1M context, and tool use.

GPT-6 Luna
10 +
GPT-6 Luna is OpenAI's budget tier in the GPT-6 family, recently price-cut to $0.10/$0.50 per million tokens.

GPT-6 Sol
160 +
OpenAI's GPT-6 mid tier — successor to GPT-5.6 Sol, at OpenAI's current post-price-cut rate.

GPT-6 Astra
800 +
OpenAI's GPT-6 flagship model, a new generation above the GPT-5.6 family.
Grok 4.7
100 +
xAI's current flagship coding and knowledge model, GA September 2026, with text + image input and a 500K token context window.

Claude Fable 5.1
200 +
Anthropic's most capable widely released model — successor to Claude Fable 5, for the most demanding reasoning and long-horizon agentic work, with a 1M token c…

Claude Opus 5.5
80 +
Anthropic's Opus 5.5 — beats Fable 5.1 on agentic benchmarks at 60% lower API price, with a 1M token context window by default.
MiniMax H3 Max Lip Sync
2000 +
MiniMax H3 Max Lip Sync turns one photo and one audio clip into a lip-synced talking video with natural expression and head movement, in any language. The outp…
Image to video
Fast generation
Lip sync

GPT Image 2.5 Sunburst
25 +
OpenAI's precision-tier GPT Image 2.5 model: heavier than Flare, tighter control across edits and subject preservation for premium creative/production workflow…
High detail
Renders text

GPT Image 2.5 Flare
25 +
OpenAI's fastest GPT Image 2.5 tier: higher quality than GPT Image 2 at 50% lower latency, the default for everyday high-quality generation. Text+image input, …
High detail
Renders text

Meta Muse Spark 1.3
70 +
Meta's Muse Spark 1.3: multimodal reasoning model tuned for longer-horizon agentic, multi-agent and coding workflows, ~20% fewer tool calls than 1.2. Omni-moda…
Long context
Wan 3.0
200 +
(50% off until Aug 30!) - Wan 3.0 generates video from a text prompt or a starting image, with cinematic motion and support for 480p, 720p, and 1080p output up…
9:16 vertical
Any aspect ratio
Supports references
Gemini 3.8 Flash
60 +
Google's most intelligent Flash-tier model, GA 2026-09-02: significant gains over 3.7 Flash on software engineering, agentic tasks and multi-step reasoning. 1.…
Long context

Tripo P1 Image to 3D
1600 +
Tripo P1 image-to-3D: one image → clean, engine-ready low-poly mesh with stable topology in seconds. Ideal for game props, web viewers and real-time pipelines.
High quality
Fast generation
Supports references

Tripo P1 Text to 3D
1600 +
Tripo P1: native 3D diffusion that generates clean, engine-ready low-poly meshes with stable topology in seconds — built for Unity/Unreal, web and real-time us…
High quality
Fast generation

Tripo H3.1 Multiview to 3D
800 +
High-detail reconstruction from several views of the same object (2–4 images): the most accurate Tripo route for characters, products and collectibles, with op…
High quality
Supports references
Controllable

Tripo H3.1 Image to 3D
800 +
Tripo's high-detail image-to-3D: one photo or concept render → production-ready mesh with PBR textures; 'Detailed' geometry preserves cloth folds, armor edges …
High quality
Supports references
Controllable

Tripo H3.1 Text to 3D
400 +
Tripo's high-detail generation: text prompt → production-ready 3D mesh with PBR textures. 'Detailed' geometry unlocks high-poly output (up to ~2M faces) for he…
High quality
Controllable

Tripo 2.5 Multiview to 3D
800 +
Builds a 3D mesh from up to four views of the same object (front required; left, back, right optional) for far better back-side and silhouette accuracy than a …
High quality
Supports references

Tripo 2.5 Image to 3D
800 +
Turns a single image into a textured 3D mesh (GLB) with optional PBR materials, HD textures and quad topology. Tripo's proven 2.5 generation via fal.
Fast generation
Cost effective
Supports references

Mistral Large 3
25 +
Mistral AI's current flagship: 675B/41B-active MoE, Apache 2.0 open weights, 262K context, tool calling and structured outputs.
Long context

Meta Muse Image
40 +
Meta's Muse Image model — faithful instruction-following and high visual fidelity, with accurate rendering of fine details like text, plots, and QR codes.
High detail
Renders text
MiniMax Speech 2.8 HD
400 +
MiniMax Speech 2.8 HD: the studio-quality tier of MiniMax's multilingual speech, same system voice library. Served via OpenRouter.
Multilingual
MiniMax Speech 2.8 Turbo
245 +
MiniMax Speech 2.8 Turbo: fast multilingual speech with MiniMax's system voice library. Served via OpenRouter.
Multilingual

Orpheus 3B
30 +
Canopy Labs' Orpheus 3B: Llama-based open TTS with seven emotive English voices. Served via OpenRouter.

Sesame CSM 1B
30 +
Sesame CSM-1B: the open conversational speech model behind Sesame's voice companions, with conversational and read-speech presets. Served via OpenRouter.

Fish Audio S2 Pro
65 +
Fish Audio S2 Pro: expressive multilingual speech with a natural default voice. Served via OpenRouter.
Multilingual

Qwen Audio 3.0 TTS Plus
80 +
Alibaba's Qwen Audio 3.0 TTS Plus: the higher-quality tier of Qwen's bilingual speech synthesis. Served via OpenRouter.
Multilingual

Qwen Audio 3.0 TTS Flash
65 +
Alibaba's Qwen Audio 3.0 TTS Flash: fast bilingual (Chinese/English) speech synthesis. Served via OpenRouter.
Multilingual

Kokoro 82M
5 +
Kokoro-82M: the tiny open-weight TTS model with 54 voices (American/British English, Spanish, French, Hindi, Italian, Japanese, Portuguese, Chinese) at a fract…

Deepgram Aura 2
125 +
Deepgram Aura-2: 90 natural voices across English, Spanish, French, German, Italian, Dutch, Japanese and more, tuned for real-time narration. Served via OpenRo…

Voxtral Mini TTS
65 +
Mistral's Voxtral Mini TTS: 30 voice presets that combine a speaker with an emotion (neutral, happy, sad, excited, angry…) in English, British English and Fren…
Multilingual

MAI Voice 2 Flash
65 +
Microsoft MAI-Voice-2 Flash: the faster, cheaper tier of MAI-Voice-2 with the same four multilingual voices. Served via OpenRouter.
Multilingual

MAI Voice 2
90 +
Microsoft MAI-Voice-2: high-fidelity multilingual text-to-speech (EN, ES, FR, DE voices). Served via OpenRouter.
Multilingual
Grok Voice TTS 1.0
65 +
xAI's Grok Voice text-to-speech: five expressive English voices (Eve, Ara, Rex, Sal, Leo). Served via OpenRouter, billed per 1k characters.
Multilingual

Meta Muse Spark 1.2
70 +
Meta's Muse Spark 1.2 flagship: omni-modal input (text, image, video, audio, files), 1M-token context, Meta's successor to the Llama line. Served via OpenRoute…
Long context
Hunyuan Hy4 Preview
45 +
Tencent's Hunyuan Hy4 (preview): 1M-token context text model with switchable reasoning, the newest Hunyuan LLM tier. Served via OpenRouter.
Long context

Qwen3.8 Flash
10 +
Alibaba's fast, low-cost Qwen3.8 tier: multimodal input, 1M-token context, tuned for coding, agents and vision tasks. Served via OpenRouter.
Long context

Qwen3.8 Max
100 +
Alibaba's Qwen3.8 Max flagship: natively multimodal (text, image, video input), 1M-token context, deep reasoning. Served via OpenRouter.

GLM-5.3
75 +
Z.ai's flagship GLM-5.3: hybrid sparse + linear attention, 1.3M-token context, strongest GLM tier for long-horizon coding and agent work. Served via OpenRouter.
Long context
MiniMax H3 Max I2V
1000 +
MiniMax H3 Max Image-to-Video animates a still image into a 5-15 second clip with native synchronized audio, and can optionally interpolate to a supplied end f…
Image to video
Cinematic
MiniMax H3 Max
1000 +
MiniMax H3 Max is a post-trained build of MiniMax H3 that generates 5-15 second videos with native synchronized audio. It ranks #1 in human preference evaluati…
Cinematic

GLM-5.3-Flash
5 +
Zhipu AI's 320B-A18B natively multimodal MoE with a 1M token context window, supersedes the GLM-5.2 base tier.
Long context
Gemini 3.7 Flash
15 +
Google's coding/agent-focused Flash tier with a 1.05M token context window, GA since 2026-08-13.

Pika 2.2
800 +
Generate video from a text prompt with Pika 2.2 — fast, stylized text-to-video with 720p/1080p output.
Cinematic

Qwen Image 3.0 Pro
160 +
Qwen-Image-3.0-Pro generates and edits images with dense, accurate text rendering, complex multi-element layouts, and photographic detail.
9:16 vertical
Any aspect ratio
Supports references

GLM-5.2
50 +
Z.AI's open-weights flagship: beats GPT-5.5 on multiple long-horizon coding benchmarks, 1M token context, 131K output.
Long context
MiniMax Music 2.6
600 +
Generate full-length songs or instrumentals from a text prompt, with optional auto-generated lyrics
LTX-2.5 Fast
240 +
Fast video generation with text-to-video and image-to-video, portrait and landscape support, synchronized audio, and frame interpolation. Up to 20 seconds at 1…
9:16 vertical
Any aspect ratio
Supports references
Seedream 5.0 Pro
180 +
ByteDance's flagship text-to-image and image editing model, generating sharp 1K and 2K images from text or up to 10 reference images
9:16 vertical
Any aspect ratio
Grok Imagine Image 2.0
160 +
xAI's Grok Imagine Image 2.0 — text-to-image generation and editing with a quality control and output up to 2k
9:16 vertical
Any aspect ratio
Supports references

Meta Muse Glimmer 30B
20 +
Meta Muse Glimmer 30B is an Apache-2.0 open-weight 30B dense agentic model from Meta, built for long tool-call sequences, hosted on Together AI.
Long context
Gemini 3.1 Pro
195 +
Gemini 3.1 Pro is Google's current Pro-tier flagship reasoning model, GA on the Gemini API with a 1M token context window.
GPT-5.6 Luna
20 +
GPT-5.6 Luna is OpenAI's budget tier in the GPT-5.6 family, offering a 1.05M token context window at a lower price point than Sol or Terra.
Long context

Seedance 2.5
4600 +
ByteDance's flagship multimodal video model with native audio, native 30-second generation, and large multimodal reference sets.
9:16 vertical
Any aspect ratio
Multi-reference
GPT-5.6 Terra
195 +
OpenAI GPT-5.6 Terra — balanced mid-tier reasoning model between Sol (flagship) and Luna (cost-efficient)
Long context

Kimi K3
240 +
Moonshot AI's current flagship: 2.8T-parameter MoE (16-of-896 experts active), Kimi Delta Attention, native vision, 1M token context.
Hunyuan Image 3
320 +
A powerful native multimodal model for image generation (PrunaAI squeezed)
9:16 vertical
Any aspect ratio
Gemini 3.5 Flash-Lite
10 +
Google's fastest, cheapest current Flash tier (~350 tok/s) — GA July 2026, for high-volume low-latency text tasks.
Gemini 3.6 Flash
15 +
Google's current Flash-tier model, GA July 2026 — successor to Gemini 3.5 Flash with faster throughput and the same agentic/coding focus.

Claude Opus 5
100 +
Anthropic's newest Opus-tier flagship, released July 2026 — successor to Opus 4.8 with the same pricing tier, built for long-horizon agentic work, knowledge ta…

Kimi K2.6
65 +
Moonshot AI's 1T-parameter MoE flagship (32B active), 256K context
Grok 4.5
100 +
xAI's current flagship model, trained with Cursor

Luma Uni-1 Max
415
Luma Uni-1 Max — higher-quality tier than Uni-1, 2K resolution.
High detail

Luma Uni-1
175
Luma autoregressive image model, matches/beats GPT Image and Nano Banana quality at lower cost. 2K resolution.
High detail

Qwen-Image-2 Pro
300 +
The pro version of Qwen Image 2 from Alibaba's Qwen team. Enhanced text rendering, realism, and semantic adherence for high-quality image generation and editin…
9:16 vertical
Any aspect ratio
Supports references
PixVerse V6
1000 +
PixVerse's flagship video generation model. Generate cinematic videos with synchronized audio, multi-shot sequences, and precise camera control.
9:16 vertical
Any aspect ratio
Supports references

Vidu Q3 Pro
280 +
High-fidelity video generation with text-to-video, image-to-video, and start-end-to-video modes. Up to 16 seconds at 1080p with synchronized audio.
9:16 vertical
Any aspect ratio
Supports references
DeepSeek V4 Flash
5 +
DeepSeek's fast, low-cost flagship model with 1M token context and dual thinking/non-thinking modes
Long context

Claude Sonnet 5
160 +
Anthropic's Claude Sonnet 5 — near-Opus quality on coding and agentic work at Sonnet-tier speed and cost, with adaptive thinking.
Lyria 3 Pro
320
Generate full-length songs up to 3 minutes from text prompts or images with Lyria 3 Pro, Google's most capable music generation model

happyhorse-1.1
1685 +
Alibaba's Happy Horse 1.1 generates videos from text, animates a single image, or builds a video from multiple reference images. Supports 720p and 1080p, 3-15 …
9:16 vertical
Any aspect ratio

Krea 2
240 +
Krea's flagship foundation image model. Larger and more flexible than Krea 2 Medium, with particular strength in photorealism and expressive artistic styles.
9:16 vertical
Any aspect ratio
Z-Image Turbo
20 +
Z-Image Turbo is a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.
GPT-5.6 Sol
160 +
OpenAI's GPT-5.6 flagship reasoning model
Long context
Grok 4.3
40 +
xAI's reasoning flagship model, successor to Grok 4
DeepSeek V4 Pro
15 +
DeepSeek's flagship reasoning model with 1M token context
Gemini 3.5 Flash
40 +
Google's most intelligent Flash-tier model for sustained frontier performance on agentic and coding tasks. Fully GA as of May 2026, succeeding Gemini 2.5 Flash.

Claude Haiku 4.5
20 +
Anthropic's fastest and most cost-effective current model — ideal for simple, speed-critical tasks. The platform's first Anthropic fast/cheap tier.

Claude Fable 5
200 +
Anthropic's most capable widely released model — for the most demanding reasoning and long-horizon agentic work. A new tier above Opus, with a 1M token context…

Claude Opus 4.8
100 +
Anthropic's most capable Opus-tier model — highly autonomous, state-of-the-art on long-horizon agentic work, knowledge work, and memory, with clearer, warmer w…

happyhorse-1.0
1685 +
Alibaba's Happy Horse 1.0 generates videos from text prompts or animates a single image into video. Supports 720p and 1080p, 3-15 second durations, and five as…
9:16 vertical
Any aspect ratio
Supports references

ltx-2.3-pro
1920 +
High-fidelity video generation with portrait support, audio-to-video, retake, and extend. Text, image, and audio-driven creation up to 4K at 50 FPS.
9:16 vertical
Any aspect ratio
Supports references

Google Veo 3.1 Lite
800 +
Google's cost-efficient video generation model with native audio, optimized for high-volume applications
9:16 vertical
Any aspect ratio
Supports references
Lyria 3
160
Generate 30-second music clips from text prompts or images with Lyria 3, Google's music generation model
Ideogram 4.0 Quality
400
The highest quality Ideogram v4 model. v4 creates images with stunning realism, creative designs, and consistent styles
GPT-5.4
75 +
OpenAI GPT-5.4 — latest generation language model with improved reasoning and creativity
High accuracy
Multi-modal
Support file upload
FLUX 2 Klein 4B
5 +
Very fast image generation. 4-step distilled, sub-second inference for production and real-time applications.
High quality
Photorealistic
Support file upload
FLUX 2 Flex
240 +
Max-quality image generation and editing with typography and up to 10 reference images.
High quality
Photorealistic
Support file upload
FLUX 2 Max
160 +
The highest fidelity image model from Black Forest Labs. Superior prompt understanding, editing consistency, and multi-reference support.
High quality
Photorealistic
Support file upload

GPT Image 2
35 +
OpenAI GPT Image 2 — state-of-the-art image generation and editing model. Highest performance with flexible image sizes, high-fidelity inputs, and precise iter…
High quality
Photorealistic
Support file upload
PixVerse v5
1405 +
PixVerse v5 — latest PixVerse with lifelike physics and striking visuals.
High quality
Image to video
Support file upload
PixVerse v4.5
1200 +
PixVerse v4.5 — upgraded version with improved motion quality and prompt adherence.
High quality
Image to video
Support file upload
PixVerse v4
1000 +
PixVerse v4 — video generation model supporting text-to-video and image-to-video up to 1080p.
High quality
Image to video
Support file upload
Wan 2.1 I2V 480p
1800 +
WaveSpeed AI serving Wan 2.1 image-to-video at 480p. Cheaper draft-tier.
High quality
Image to video
Support file upload
Wan 2.1 I2V 720p
5000 +
WaveSpeed AI serving Wan 2.1 image-to-video at 720p.
High quality
Image to video
Support file upload
Wan 2.1 T2V 720p
4800 +
WaveSpeed AI serving Wan 2.1 text-to-video at 720p.
High quality
Image to video
9:16 vertical
Wan 2.5 T2V Fast
1360 +
Alibaba Wan 2.5 T2V Fast — fast variant of Wan 2.5 text-to-video for quick iteration.
High quality
Image to video
Fast generation
Wan 2.2 T2V Fast
200 +
Alibaba Wan 2.2 T2V Fast — fast text-to-video at 480p. Cheap and quick for iteration.
High quality
Image to video
Fast generation
OpenAI Sora 2 Pro
4800 +
OpenAI Sora 2 Pro — highest-fidelity Sora 2 tier for premium cinematic video. Text-to-video and image-to-video.
Image to video
Cinematic
Support file upload
OpenAI Sora 2
1600 +
OpenAI Sora 2 — flagship video model with cinematic quality, strong physics simulation, and precise prompt adherence. 4-12s at 720p/1080p.
Image to video
Cinematic
Support file upload
Kling 2.6
1405 +
Kling 2.6 — text-to-video and image-to-video with native audio. 5-10s with motion prompts and negative prompts.
High quality
Image to video
Support file upload
Kling 2.5 Turbo Pro
1405 +
Kling 2.5 Turbo Pro — image-to-video with cinematic motion and precise intent following. Up to 10s.
Image to video
Fast generation
Support file upload
Pruna P-Video
20 +
Pruna AI's multimodal video model built for speed and iteration — supports text, image, and audio-to-video in a unified endpoint. 1-10s, up to 1080p.
High quality
Image to video
Support file upload
Wan 2.7 VideoEdit
800 +
Alibaba Wan 2.7 VideoEdit — instruction-based video editing model. Modify an existing clip with natural-language instructions: background swaps, lighting, styl…
High quality
Image to video
Support file upload

GPT Image 1.5
40 +
OpenAI GPT Image 1.5 — latest flagship image gen/edit model with stronger preservation, precise iterative edits, up to 4x faster than GPT Image 1. Built on GPT…
High quality
Photorealistic
Any aspect ratio
Ideogram v3 Quality
360 +
Ideogram v3 Quality — highest fidelity variant of Ideogram v3 for premium graphic design and typography work.
Renders text
Best for logos
Support file upload
Ideogram v3 Turbo
120 +
Ideogram v3 Turbo — fast, design-focused image generation with best-in-class typography, accurate text rendering and graphic design quality at a low price.
Renders text
Best for logos
Support file upload
Seedream 5 Lite
140 +
ByteDance Seedream 5 Lite — lightweight variant of Seedream 5 with built-in multi-step reasoning, example-based editing, and deep domain knowledge. Up to 3K.
High quality
Photorealistic
Support file upload
Seedream 4.5
160 +
ByteDance Seedream 4.5 — upgraded unified image generation and editing model with stronger spatial understanding, world knowledge, and cinematic visuals. Up to…
High quality
Photorealistic
Support file upload
Hailuo 2.3 Pro I2V
1960 +
MiniMax Hailuo 2.3 Pro image-to-video (1080p) on fal.ai — flagship tier for quality.
Image to video
Fast generation
Support file upload
Kling 3 Pro I2V
3365 +
Kling Video v3 Pro image-to-video on fal.ai — cinematic visuals, fluid motion, native audio, element referencing. Up to 15s.
High quality
Image to video
Support file upload
Seedance 2.0 Fast
1000 +
ByteDance Seedance 2.0 Fast — speed-optimized variant of Seedance 2.0 with native audio-video joint generation. Supports text-to-video and image-to-video at 48…
High quality
Image to video
Cinematic
MiniMax Hailuo 2.3 Fast
760 +
MiniMax Hailuo 2.3 Fast — faster and cheaper variant of Hailuo 2.3 for quick iterations.
High quality
Fast generation
Support file upload
MiniMax Hailuo 2.3
1120 +
MiniMax Hailuo 2.3 — latest flagship video model with improved motion, physics, and prompt adherence. Supports text-to-video and image-to-video up to 1080p.
High quality
Fast generation
Support file upload
Wan 2.7 T2V
2000 +
Alibaba Wan 2.7 — newest 27B open-weights video model. Generates up to 15s 1080p with native audio-sync from text prompts.
High quality
Image to video
9:16 vertical
Google Veo 3.1 Fast
2000 +
Google Veo 3.1 Fast — faster, cheaper variant of Veo 3.1 with native audio. Great for high-volume production.
Image to video
Cinematic
Support file upload
Google Veo 3.1
3200 +
Google Veo 3.1 — state-of-the-art text-to-video model with native audio generation, natural lip sync, cinematic quality, and strong prompt adherence. Up to 108…
Image to video
Cinematic
Support file upload

Claude 4.7 Opus
100 +
Latest balanced AI model combining speed with intelligence—analyzes images, processes large documents, and excels at programming tasks with multilingual support
Multi-modal
Cost effective
Support file upload

Claude 4.6 Opus
100 +
Latest balanced AI model combining speed with intelligence—analyzes images, processes large documents, and excels at programming tasks with multilingual support
Multi-modal
Cost effective
Support file upload

Claude 4.6 Sonnet
60 +
Latest balanced AI model combining speed with intelligence—analyzes images, processes large documents, and excels at programming tasks with multilingual support
Multi-modal
Cost effective
Support file upload
Nano Banana 2
270 +
Google's fast image generation model with conversational editing, multi-image fusion, and character consistency
High quality
Any aspect ratio
Seedance 2.0
3600 +
ByteDance's multimodal video generation model with native audio, multimodal reference inputs, and intelligent duration control. Supports text-to-video and imag…
High quality
Image to video
Cinematic
gen-4.5
12000 +
State-of-the-art video motion quality, prompt adherence and visual fidelity
High fidelity
High-quality motion
Image to video
Kling V3 Video
2020 +
Cinematic videos up to 15 seconds with multi-shot control, native audio, and improved consistency.
High quality
Image to video
Support file upload
Kling V3 Omni Video
2020 +
Unified multimodal video generation with reference images, video editing, native audio, and multi-shot control.
High quality
Image to video
Support file upload
Kling V3 Motion Control
1405 +
Transfer motion from a reference video to any character image with improved consistency and quality.
High quality
Image to video
Support file upload
Grok Imagine Video
200 +
Generate videos using xAI Grok Imagine Video model. Fast ~30s generation with audio.
High quality
9:16 vertical
pia
140 +
Personalized Image Animator
High quality
Image to video
9:16 vertical
i2vgen-xl
30
RESEARCH/NON-COMMERCIAL USE ONLY: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models
High quality
Image to video
9:16 vertical
mochi-1
140
Mochi 1 preview is an open video generation model with high-fidelity motion and strong prompt adherence in preliminary evaluation
High quality
9:16 vertical
motion-2.0
1200 +
Create 5s 480p videos from a text prompt
High quality
9:16 vertical
controlvideo
140 +
Training-free Controllable Text-to-Video Generation
High quality
9:16 vertical
cogvideox-5b
140 +
Generate high quality videos from a prompt
High quality
Image to video
9:16 vertical
hunyuan-video
400 +
A state-of-the-art text-to-video generation model capable of creating high-quality videos with realistic motion from text descriptions
High quality
Image to video
9:16 vertical
ltx-video
140 +
LTX-Video is the first DiT-based video generation model capable of generating high-quality videos in real-time. It produces 24 FPS videos at a 768x512 resoluti…
High quality
9:16 vertical
tile-morph
30 +
Create tileable animations with seamless transitions
High quality
9:16 vertical

stable_diffusion_infinite_zoom
30 +
Use Runway's Stable-diffusion inpainting model to create an infinite loop video
High quality
9:16 vertical

stable-diffusion-animation
30 +
Animate Stable Diffusion by interpolating between two prompts
High quality
9:16 vertical
videocrafter
140
VideoCrafter2: Text-to-Video and Image-to-Video Generation and Editing
High quality
9:16 vertical
animatediff-prompt-travel
30 +
🎨AnimateDiff Prompt Travel🧭 Seamlessly Navigate and Animate Between Text-to-Image Prompts for Dynamic Visual Narratives
High quality
9:16 vertical
animatediff-illusions
30 +
Monster Labs' Controlnet QR Code Monster v2 For SD-1.5 on top of AnimateDiff Prompt Travel (Motion Module SD 1.5 v2)
High quality
9:16 vertical
animate-diff-og
30 +
Animate Your Personalized Text-to-Image Diffusion Models
High quality
9:16 vertical
animate-diff
30 +
🎨 AnimateDiff (w/ MotionLoRAs for Panning, Zooming, etc): Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
High quality
9:16 vertical
hotshot-xl
30 +
😊 Hotshot-XL is an AI text-to-GIF model trained to work alongside Stable Diffusion XL
High quality
9:16 vertical
zeroscope-v2-xl
450
Zeroscope V2 XL & 576w
High quality
9:16 vertical
text2video-zero
560
Text-to-Image Diffusion Models are Zero-Shot Video Generators
High quality
9:16 vertical
damo-text-to-video
30
Multi-stage text-to-video generation
High quality
9:16 vertical
Runway Gen-4.5
2400 +
Runway Gen-4.5 — highest quality video generation model with state-of-the-art motion and visual fidelity
High quality
Image to video
9:16 vertical
Runway Gen-4 Turbo
1000 +
Runway Gen-4 Turbo — fast, high-quality video generation from images with improved motion coherence and cinematic controls
High quality
Image to video
Fast generation
Recraft V4 Pro SVG
1200 +
Generate detailed SVG vector graphics from text prompts. Recraft V4 Pro's design taste with more geometric detail and finer paths — clean layers, editable outp…
High quality
Renders text
Best for logos
Recraft V4 SVG
320 +
Generate production-ready SVG vector images from text prompts. Recraft V4's design taste applied to vector output — clean geometry, structured layers, and edit…
Renders text
Best for logos
Any aspect ratio
Recraft V4 Pro
1000 +
Recraft's latest image generation model at ~2048px resolution. Same design taste and prompt accuracy as V4, with higher resolution for print-ready and large-sc…
High quality
Renders text
Best for logos
Recraft V4
160 +
Recraft's latest image generation model, built around design taste. Strong prompt accuracy, art-directed composition, and integrated text rendering. Fast and c…
Renders text
Best for logos
Any aspect ratio
p-image-edit
40 +
P-Image Edit by Pruna AI edits reference images using text instructions. Generation starts at 40 coins. Review the quote before sending; settings can change th…
Editing presets
Multi-reference
Sub-second speed
p-image
20 +
P-Image by Pruna AI creates images from text prompts. Generation starts at 20 coins. Review the quote before sending; settings can change the price. Supports m…
Custom sizes & aspect
Optional safety off
Sub-second generation
crystal-upscaler
200 +
Crystal Upscaler is a high-precision AI image upscaler optimized for portraits, faces, and product photography, powered by Clarity AI technology. This speciali…
Adjustable creativity
Best for faces
Official Replicate model
seedance-1.5-pro
105 +
SeeDANCE 1.5 Pro by ByteDance is a revolutionary joint audio-video AI model that generates cinematic videos with synchronized audio from text descriptions or i…
Duration, FPS & aspect
Optional synced audio
Text & image to video
FLUX 2 Pro
60 +
FLUX 2 Pro by Black Forest Labs is a professional-grade AI image generator designed for brand consistency and creative control. This advanced text-to-image and…
8 reference images
Any aspect ratio
Creative control
Wan 2.5 T2V
1000 +
Wan 2.5 T2V (Text-to-Video) by Alibaba is an advanced AI video generator that creates cinematic videos from text descriptions with optional audio synchronizati…
Audio sync
Fast processing
Flexible resolution
Wan 2.5 Image to Video
1000 +
Wan 2.5 Image to Video (I2V) by Alibaba is a powerful AI animator that transforms static images into cinematic videos with optional background audio synchroniz…
Audio sync
Custom duration
Flexible resolution
Seedream 4
120 +
Seedream 4 by ByteDance is a unified AI image generation and editing model capable of creating stunning images up to 4K resolution (4096×4096 pixels) from text…
4K output
Accurate editing
High detail
Nano Banana Pro
140 +
Nano Banana Pro by Google is the enhanced version of Google's experimental AI image generator and editor, offering improved quality and advanced natural langua…
Advanced image editing
Fast processing
Multi-image
Nano Banana
160 +
Nano Banana by Google is an experimental AI image generator and editor that leverages natural language processing for intuitive visual creation and transformat…
Advanced image editing
Fast processing
Multi-image
Seedance 1 Lite
290 +
Seedance 1 Lite by ByteDance is an affordable, high-speed AI video generator that creates professional cinematic videos from text prompts or static images with…
5–10 sec clips
Any aspect ratio
Flexible resolution
Pyramid Flow
200 +
Generate short high-quality videos from text or images quickly—budget-friendly option for social media content and rapid prototyping
High quality
Multi-modal
Cost effective
SDXL Pixar
15 +
Generate Pixar-style poster art from text or image inputs
Customizable
High quality
Multi-modal
Omni-Human
500 +
Omni-Human by ByteDance is a revolutionary AI video generator that creates highly realistic lip-synced talking head videos from a single static photograph and …
High quality
High accuracy
Multi-modal

GPT-OSS 120B
5 +
GPT-OSS 120B is OpenAI's open-weight 120-billion parameter language model designed for customization, on-premise deployment, and full enterprise control. This …
Creative
Efficient
Fast processing
Wan 2.2 I2V Fast
200 +
Generate cinematic videos from images with fast, accurate control
High quality
Fast generation
Multi-modal
SeeDANCE 1 Pro Fast
120 +
Generate cinematic 5–10s 1080p videos from text or images
Supports references
High quality
Fast generation
SeeDANCE 1 Pro
240 +
Create cinematic 5-10 second 1080p videos from text or images—ByteDance's professional video generator for ads, social content, and previews
Supports references
High quality
Fast generation
ACE-Step Audio
10 +
ACE-Step Audio is an advanced AI music generator that transforms text prompts into professional-quality audio tracks. This text-to-music AI model excels at cre…
Fine control
High quality
Fast generation
Recraft V3
160 +
Create print-ready designs with flawless text, precise layout, and vectors
Vector output
High quality
Large context
Clarity Upscaler
40 +
AI image upscaler that improves resolution, clarity, and style
Creative control
High quality
Open source
Video Upscaler
120 +
AI video upscaler enhancing low-resolution footage to 1080p/4K/8K with detail reconstruction—restore old videos and enhance quality for modern displays
Fast processing
Noise reduction
High quality

LLama 3.3 70B
5 +
Llama 3.3 70B Instruct Turbo by Meta is a powerful open-source instruction-tuned AI language model with 70 billion parameters, optimized for extended context u…
Instruction-tuned
High accuracy
Multilingual
Flux Pro Fill
200
Flux Pro Fill by Black Forest Labs is a professional AI inpainting and outpainting tool for seamless image expansion, object removal, and completion of partial…
High quality
Fast generation
Multi-modal
QR Code Generator
5 +
Generate branded, secure QR codes with dynamic, trackable designs
Customizable
Real-time analytics
High accuracy
Flux Kontext Max
320 +
Flux Kontext Max by Black Forest Labs is an advanced multi-scene storytelling AI image generator that creates coherent visual narratives spanning multiple conn…
High consistency
Fast generation
Multi-modal
Flux Kontext Pro
160 +
Flux Kontext Pro by Black Forest Labs is an iterative AI image editor that generates and progressively refines visuals through multiple revision cycles, enabli…
High consistency
Fast generation
Multi-modal
Flux Pro Canny
200
Precision retexturing tool maintaining original structure while applying new styles—transform images based on edge detection for architectural visualization an…
Supports references
Flux Pro 1.1 Redux
200
Advanced image-to-image transformer specializing in style transfer and artistic remixing—reimagine photos with different aesthetics, lighting, and artistic sty…
High quality
Fast generation
Supports references
Flux Pro Ultra 1.1
240 +
Flux Pro Ultra 1.1 by Black Forest Labs is an ultra-high-resolution photorealistic AI image generator creating stunning 4-megapixel (4MP) outputs optimized for…
Versatile modes
High quality
High accuracy
MiniMax (Hailuo AI)
2000 +
Open-source video model by MiniMax generating high-quality cinematic videos—free alternative for filmmakers and content creators needing professional results
Supports references
Mochi v1
1600 +
Text-to-video AI creating high-fidelity realistic motion and scenes—excellent for storytelling, advertising, and creative video projects
Customizable
High quality
High accuracy

Luma Ray 2 Flash
2200 +
Fast photorealistic video creation from text or images with rapid turnaround—perfect for quick content production and social media campaigns
Cinematic control
High quality
Multi-modal

Luma Ray 2
6400 +
Luma Ray 2 represents the next generation of AI video synthesis, delivering photorealistic, professional-grade video generation from text descriptions or image…
Cinematic control
High quality
Multi-modal

Stable Diffusion 3.5 Large
260 +
High-quality text-to-image and image-to-image at 1MP, strong prompt adherence
High quality
High accuracy
Multi-modal
Flux Pro 1.1
160
Flux Pro 1.1 by Black Forest Labs is a high-speed professional AI image generator producing 2K resolution outputs with exceptional prompt accuracy and rapid ge…
High quality
High accuracy
Fast generation
Virtual Try On
160 +
Virtual Try On — realistic apparel & jewelry previews on you
Cross-platform
Real-time
High accuracy
Ideogram Upscaler
240 +
AI upscaler doubling image resolution with enhanced detail recovery and intelligent cropping—sharpen low-res images for print and high-DPI displays
Easy integration
High quality
Fast generation
Ideogram v2
320 +
Advanced text-to-image model with photorealistic quality and superior typography—generate professional visuals with embedded text for branding and design
Custom styles
High quality
High accuracy
Ideogram v2 Turbo
200 +
High-fidelity image generation with flexible style control and fast output—create diverse visual content from photorealistic to artistic with text integration
High quality
Fast generation
Multi-modal

Sonar
5 +
Fast factual search & reasoning; also multilingual multimodal embeddings
Fast
High accuracy
Multi-modal

Sonar Pro
60 +
Real-time web search and synthesis, fast, cited answers
Real-time
Supports references
Fast generation
Gemini 2.5 Flash
10 +
Efficient multimodal AI understanding images, audio, and video at high speed—cost-effective solution for content moderation, media analysis, and automated tran…
Low latency
Multi-modal
Cost effective
Gemini 2.5 Pro
40 +
Google's most capable AI processing massive multimodal datasets with advanced reasoning—ideal for academic research, complex data analysis, and scientific comp…
Fast response
High accuracy
Multi-modal
ToonCrafter
80
Cartoon animation generator creating smooth transitions between keyframe images—produce animated shorts, explainer videos, and motion comics
User-friendly
High quality
High accuracy
Video Morpher
80
Blend multiple images with seamless morphing transitions—create mesmerizing visual effects for music videos, presentations, and artistic projects
Supports references
Face Swap
5
Seamlessly swap faces in photos and videos with photoreal results
Supports video
High quality
High accuracy
Blend Images
150
Blend Images is an AI-powered image compositing tool that seamlessly merges multiple photos into realistic, cohesive compositions. This image-to-image model in…
High quality
Fast generation
Supports references
Kandinskiy 2.2
60 +
Generate photorealistic images from text, edit and blend images
Fine control
High quality
BG Remover
Free
BG Remover is a powerful AI background removal tool that automatically isolates subjects from their backgrounds with pixel-perfect precision. This image-to-ima…
Easy integration
Fast processing
High accuracy
Instant ID
150
Zero-shot identity-preserving image generation from one face
High accuracy
Fast generation
Supports references

Latent Consistency
5
Generate high-quality images from text in under a second
Customizable
High quality
Fast generation

o4 mini
20 +
Fast multimodal reasoning AI understanding images and text—combines visual analysis with mathematical logic for data science, engineering diagrams, and technic…
High accuracy
Fast generation
Multi-modal

o3
35 +
OpenAI's reasoning powerhouse with multi-step thought processes—solves advanced mathematics, writes complex algorithms, and performs deep scientific analysis
Advanced reasoning
Tool use
High accuracy

o3 mini
20 +
Compact reasoning model optimized for STEM tasks—delivers accurate solutions in mathematics, physics, chemistry, and programming at lower cost
Fast response
High accuracy
Multi-modal

MusicGen Remixer
500 +
AI music remixer with chord-aware controls for customizing generated tracks—remix, rearrange, and fine-tune AI music with harmonic precision
High quality
Fast generation
Supports references

Bark
50
Bark by Suno AI is a revolutionary multilingual text-to-speech model capable of generating highly realistic speech, music, and sound effects across 100+ langua…
High quality
Fast generation
Multilingual
Sticker Maker
120 +
Generate stickers from text or photos — fast, editable, high‑res
High quality
Fast generation
Multi-modal
Instruct pix2pix
30 +
Text-guided image editor — fast, precise image-to-image edits
Fast inference
High accuracy
Multi-modal

SDXL Realism 2.0
50 +
Generate photorealistic images and portraits with cinematic lighting
Image input
High quality
High accuracy
Style transfer
50
Create images in style of uploaded image
High quality
Fast generation
Supports references

Stable Diffusion 3 Turbo
160 +
Fast text-to-image & image-to-image generation, excellent typography
Typography
High quality
Fast generation
Face to many
80 +
Identify individuals by matching one face against millions
Fast matching
Scalable
High accuracy
Image Upscaler
10
Upscale and restore images with AI for sharper, print-ready results
Fast processing
Scalable
User-friendly

Llama 3 8B
5 +
A powerful, conversational AI model optimized for natural language understanding and generation tasks.
High quality
Open source
Multilingual

Stable Audio
40 +
Stability AI's music and sound generator from text or audio prompts—create professional audio tracks, ambiences, and soundscapes for creative projects
High quality
Large context

Stable Diffusion 3
260 +
Generate high-resolution images from text and images, fast and customizable
High quality
Fast generation
Multi-modal

Stable Diffusion 3 Medium
140 +
Generate photorealistic images from text; runs on consumer hardware
High quality
Fast generation
Multi-modal

Stable Diffusion Core
120 +
Generate detailed images from text; inpainting, outpainting, edits
High quality
Fast generation

SDXL Flash
20
Fast sdxl with higher quality
Diverse styles
High speed
Text to image

Claude 4.5 Opus
100 +
Latest balanced AI model combining speed with intelligence—analyzes images, processes large documents, and excels at programming tasks with multilingual support
Multi-modal
Cost effective
Support file upload

Claude 4.5 Sonnet
60 +
Latest balanced AI model combining speed with intelligence—analyzes images, processes large documents, and excels at programming tasks with multilingual support
Multi-modal
Cost effective
Support file upload

Claude 4.1 Opus
1200 +
Enhanced AI assistant with extended memory for long-term projects—maintains context across sessions for ongoing coding work, research, and collaborative writing
Agentic tools
Safe alignment
High accuracy
GPT-5.4
240 +
OpenAI GPT-5.4 — latest generation language model with improved reasoning and creativity
High accuracy
Multi-modal
Support file upload
GPT-5.2
225 +
Latest iteration with breakthrough reasoning abilities and multimodal understanding—excels at complex problem-solving, advanced mathematics, and enterprise-lev…
High accuracy
Multi-modal
Support file upload
GPT-5-mini
35 +
Efficient multimodal AI processing images and large documents at high speed—optimized for rapid content generation, summarization, and real-time applications
Fast generation
Cost effective
Large context
GPT-5
160 +
GPT-5 represents OpenAI's next-generation flagship AI model with breakthrough capabilities in advanced reasoning, multimodal understanding, and sophisticated c…
High accuracy
Multi-modal
Support file upload
Grok 4
60 +
Understands text, images and voice; excels at coding and reasoning
High accuracy
Multi-modal
Support file upload

GPT-4.1
35 +
Advanced multimodal AI understanding text, images, video, and large files with superior coding capabilities—excellent for technical documentation, data analysi…
Multi-modal
Cost effective
Support file upload
Deepseek
15 +
Generates text, understands images and code; excels at reasoning
Fast reasoning
High accuracy
Multi-modal

GPT-4o
40 +
GPT-4o (GPT-4 Omni) by OpenAI is a lightning-fast, cost-efficient multimodal AI model that processes both text and images with exceptional contextual understan…
Support file upload

MusicGen
5 +
MusicGen by Meta AI is a sophisticated music generation model that creates high-quality, original music tracks from text descriptions or audio references. This…
High quality
Supports references
Controllable

Stable Diffusion XL
10 +
Stable Diffusion XL (SDXL) by Stability AI is a powerful open-source text-to-image and image-to-image AI model capable of generating ultra-high-resolution 1024…
High quality
Fast generation
Multi-modal

GPT-4o Mini
5 +
GPT-4o Mini is OpenAI's most cost-effective multimodal AI model, offering an optimal balance between performance and affordability. This compact yet powerful m…
Fast generation
Multi-modal
Cost effective
Flux Schnell
15 +
Flux Schnell (German for 'fast') by Black Forest Labs is an ultra-high-speed AI image generator optimized for instant visual creation. This lightning-fast text…
High quality
Fast generation
Cost effective