Captured photo
MiniMax H3 Max Lip Sync
2000 +

Percs

Image to video
Fast generation
Lip sync
Support file upload
Multilingual

About

MiniMax H3 Max Lip Sync turns one photo and one audio clip into a lip-synced talking video with natural expression and head movement, in any language. The output matches the audio length (5–14.8 s) and the aspect ratio of the photo. Ranked #1 for quality and speed in fal's lip-sync evals, with a median generation time of ~11 seconds.

Settings

Resolution-  Output resolution. Price scales with resolution.
Transcription guidance-  Transcribe the audio to guide lip sync. Helps on clear speech; leave off for singing or noisy audio.
Seed-  Random seed for reproducible results. Leave empty for random.