Read The Day

AIReality checkStory 06

Vision models keep choosing the same experiment

What changedA controlled benchmark changed the physical question and the cheapest useful follow-up measurement, yet six open vision-language models repeated the same experimental choice on 95.1% to 100% of paired cases. The finding separates answering from evidence gathering: a model can sound physically fluent while failing to notice which new measurement would actually resolve uncertainty.

Diagram showing how changed physical evidence should lead to a different follow-up experiment

The useful part

Why it matters

Embodied systems need to decide what evidence to collect next, so final-answer benchmarks can hide a consequential planning weakness.

Worth doing

What to do next

Add paired tests where the optimal observation changes before trusting a model to choose sensors or experiments autonomously.

Keep in mind

Good to know

This is a small controlled benchmark, not a measurement of closed-loop robot or laboratory performance.

Evidence

Primary source

Experiment Selection authors

Read the complete 13 September 2026 edition