Read The Day

AIAudio generationStory 02

One model learns speech, music and sound

What changedStepFun described one autoregressive model for speech, designed voices, vocals, effects, music and mixtures through a shared audio token space. A general audio model could replace several separate specialist systems. The result is concrete, but it remains research evidence rather than a production guarantee.

Human-likeness evaluation chart for StepAudio 3 Gen speech generation

The useful part

Why it matters

A general audio model could replace several separate specialist systems.

Worth doing

What to do next

Reproduce the core result against your own data, hardware and failure cases before depending on it.

Keep in mind

Good to know

Benchmarks are author-reported. Production reliability is unproven.

Evidence

Primary source

StepAudio 3 Gen authors

Read the complete 14 September 2026 edition