AIAgent evaluationStory 05
Agents still stumble after one tool goes wrong
What changedParaRecover adds 10,626 multi-turn cases across 14 error types, then scores whether agents preserve structure, diagnose the fault and choose a useful recovery strategy. An agent that finishes easy runs can still fail badly when one tool call poisons several dependent branches; this benchmark makes that weakness visible. The result is concrete, but it remains research evidence rather than a production guarantee.

The useful part
Why it matters
An agent that finishes easy runs can still fail badly when one tool call poisons several dependent branches; this benchmark makes that weakness visible.
Worth doing
What to do next
Reproduce the core result against your own data, hardware and failure cases before depending on it.
Keep in mind
Good to know
Benchmark performance may not predict recovery in every production tool stack. Model results are author-reported and the rubric is newly introduced.
Evidence