Drop a neutral face photo and a driving clip. Fit the four points on each, then Reenact — the photo's whole face (mouth, brows, eyes) follows the clip.
Scale is normalized by the head↔chin distance you set on each face (not auto guesswork), so mismatched aspect ratios line up. All 468 landmarks warp — mouth, brows, and eyes/lids. Next: drive this from VIAxVOICE audio (no clip) so a still photo speaks any line.