Descript Adds Generative Lip-Sync to Its AI Video Translation and Dubbing Feature
Image via provideocoalition.com
Descript's video translation and dubbing tool now goes beyond traditional rotoscoping by using generative AI to reconstruct the lower half of a speaker's face to match the mouth movements, pacing, and sounds of a newly dubbed language. The system encodes the original video into a latent space, regenerates the mouth region using the translated audio while preserving lighting and speaker identity, and then blends the result seamlessly with the untouched portions of the frame.
A hands-on review tested the feature across multiple languages — two varieties of French, German, and Greek — and found the lip-sync quality notably accurate. The reviewer, a certified translator, notes that while the AI-only translation is usable for quality assessment, curated human-native-speaker review (available on Descript's Business and Enterprise plans) remains preferable for production-ready multilingual content.