Overview
Lip Sync synchronizes lip movements in videos with new audio tracks using advanced AI motion analysis. The API creates realistic lip-sync animation that matches speech patterns, timing, and mouth movements for natural-looking results. Typical processing time: ~3m 49s median (see all processing times)How It Works
- Provide a source video - Upload the video with the person speaking
- Provide audio - Upload the new audio track to sync to
- API processes the video - AI analyzes audio and animates lip movements
- Download the result - Retrieve your lip-synced video
Use Cases
- Dubbing and localization - Translate videos to new languages with matching lips
- Personalized messages - Create custom video messages with any voice
- Educational content - Produce training videos with voiceovers
- Entertainment - Create fun lip-sync content for social media
- Accessibility - Add voiceovers to silent video content
Best Practices
Video Requirements
- Face visibility - Full face visible with minimal obstructions
- Good lighting - Even lighting on the face
- Stable framing - Face stays in frame throughout
- Moderate motion - Avoid extreme head movements
Audio Requirements
- Clear speech - Well-recorded audio without background noise
- Appropriate length - Audio duration determines output length
- Supported formats - MP3, WAV, AAC, FLAC, M4A, OPUS, OGG/OGA, WEBM/WEBA, AIFF, AMR
- Natural pacing - Normal speaking pace for best results
Matching Audio to Video
Code Examples
Basic Lip Sync
Pricing
Lip Sync is charged per rendered frame, and the rate depends onstyle.generation_mode:
For example, a 10-second clip capped at 30 FPS in
lite mode costs about 300 credits. Use max_fps_limit to lower the frame rate and reduce cost. Credits are only charged for the frames that actually render, and the completed job’s credits_charged shows the exact cost.
API Reference
Lip Sync API Reference
View full API specification
Related Tools
AI Voice Generator
Generate audio with celebrity voices
Face Swap Video
Replace faces in videos