New Audio, Matching Mouth
The lip sync API takes a video of a speaker and a separate audio track and returns the video with the speaker’s mouth movement regenerated to match the audio. The rest of the frame is untouched. It is the endpoint behind dubbing into another language, replacing a line without a reshoot, and turning a generated avatar into a talking head.
It does not care where the audio came from. Record it, generate it with the text-to-speech endpoint, or generate it in a voice you cloned from a 4-second sample — all three run on the same ModelsLab key, so a full dubbing pipeline is two requests.