MiniMax H3 Max Lip Sync
2000 +
Percs
Image to video
Fast generation
Lip sync
Support file upload
Multilingual
About
MiniMax H3 Max Lip Sync turns one photo and one audio clip into a lip-synced talking video with natural expression and head movement, in any language. The output matches the audio length (5–14.8 s) and the aspect ratio of the photo. Ranked #1 for quality and speed in fal's lip-sync evals, with a median generation time of ~11 seconds.
Settings
Resolution- Output resolution. Price scales with resolution.
Transcription guidance- Transcribe the audio to guide lip sync. Helps on clear speech; leave off for singing or noisy audio.
Seed- Random seed for reproducible results. Leave empty for random.