AIGenerated videoStory 05
Talking avatars stream with smaller models
What changedA new talking-head system generates motion in a compact causal space, then distils both that motion model and the video renderer from a frozen teacher. The authors report 15.4 frames per second with 1.3 seconds of latency. That is not instant conversation, but it moves high-quality avatar video closer to a live stream without requiring the full teacher at runtime.

The useful part
Why it matters
Smaller streaming systems could make responsive avatars cheaper to run, while the measured delay gives product teams a concrete usability constraint.
Worth doing
What to do next
Test turn-taking and lip synchronization under real network conditions before treating the laboratory frame rate as an interactive experience.
Keep in mind
Good to know
Quality and latency are reported by the authors, and 1.3 seconds can still feel slow in rapid conversation.
Evidence