RoboticsRobot safetyStory 05
Long conversations can make a robot ignore rules
What changedA new benchmark ran language-model robot controllers through forty conversations of one hundred turns each, looking for rule drift and unsafe decisions. The published first layer shows that long dialogue can erode instruction following. Physical validation is still under way, so the current result is a warning signal rather than a complete safety verdict.

The useful part
Why it matters
A robot’s high-level language model can become either too cautious or dangerously permissive as a conversation grows, even when low-level rules look fixed.
Keep in mind
Good to know
Published results are primarily text-layer tests; physical Unitree validation remains ongoing. The benchmark is a preprint and does not represent all robot safety architectures.
Evidence