These sentences were never performed — the face has no recording of them. It mouths them live from a viseme table learned from your capture: the envelope says when the mouth opens, the words say what shape it makes. This is the generalization — any line VIAxVOICE speaks, the face can mouth.
Engine: text → viseme (AH open · EE spread · OO round · M/B/P closed · F/V) × the envelope for timing, driving mouth shape targets learned from your performance. Expression (blinks, brow) idles for now — the next layer is emotion per line. Then a face per voice.