Read The Day

RoboticsRobot safetyStory 05

Long conversations can make a robot ignore rules

What changedA new benchmark ran language-model robot controllers through forty conversations of one hundred turns each, looking for rule drift and unsafe decisions. The published first layer shows that long dialogue can erode instruction following. Physical validation is still under way, so the current result is a warning signal rather than a complete safety verdict.

Top view of the physical robot environment used for long-conversation safety tests

The useful part

Why it matters

A robot’s high-level language model can become either too cautious or dangerously permissive as a conversation grows, even when low-level rules look fixed.

Keep in mind

Good to know

Published results are primarily text-layer tests; physical Unitree validation remains ongoing. The benchmark is a preprint and does not represent all robot safety architectures.

Evidence

Primary source

Aulon Bajrami, Mohamed Elshamouty and Werner Kraus

Read the complete 9 September 2026 edition