AIReality checkStory 06
Vision models keep choosing the same experiment
What changedA controlled benchmark changed the physical question and the cheapest useful follow-up measurement, yet six open vision-language models repeated the same experimental choice on 95.1% to 100% of paired cases. The finding separates answering from evidence gathering: a model can sound physically fluent while failing to notice which new measurement would actually resolve uncertainty.

The useful part
Why it matters
Embodied systems need to decide what evidence to collect next, so final-answer benchmarks can hide a consequential planning weakness.
Worth doing
What to do next
Add paired tests where the optimal observation changes before trusting a model to choose sensors or experiments autonomously.
Keep in mind
Good to know
This is a small controlled benchmark, not a measurement of closed-loop robot or laboratory performance.
Evidence