This is the hybrid: a real face's luminance as the base, deformed live by text → visemes → mouth/brow/eye motion. No driving clip — the face wears the voice.
The face's mouth, brows and lids are moved by displacing the 468-landmark mesh from the viseme signals, then the photo is triangle-warped to follow — the reverse of the capture pipeline. Next: drive the timing from real Kokoro audio (the amplitude envelope) instead of a procedural pace, so it speaks any /synthesize line in perfect sync. — Kairos