Read The Day

AIGenerated videoStory 05

Talking avatars stream with smaller models

What changedA new talking-head system generates motion in a compact causal space, then distils both that motion model and the video renderer from a frozen teacher. The authors report 15.4 frames per second with 1.3 seconds of latency. That is not instant conversation, but it moves high-quality avatar video closer to a live stream without requiring the full teacher at runtime.

Qualitative examples from the streaming talking-head generation system

The useful part

Why it matters

Smaller streaming systems could make responsive avatars cheaper to run, while the measured delay gives product teams a concrete usability constraint.

Worth doing

What to do next

Test turn-taking and lip synchronization under real network conditions before treating the laboratory frame rate as an interactive experience.

Keep in mind

Good to know

Quality and latency are reported by the authors, and 1.3 seconds can still feel slow in rapid conversation.

Evidence

Primary source

Streaming Talking Head authors

Read the complete 10 September 2026 edition