Scale Labs Benchmark Shows GPT-6-astra Scoring 53.6% on Intuitive Visual Reasoning, Humans 93.1%
Summary
Scale Labs introduces Humanity's Sixth Sense (HSS), a benchmark of image and video items testing intuitive visual reasoning across temporal, spatial, social and abstract structure. Human participants reach 93.1% accuracy, while the top model, GPT-6-astra, scores 53.6% at maximum reasoning effort. An agentic setup with dynamic visual manipulation narrows but does not close the gap, the researchers say.
Key Points
- Scale Labs researchers say the benchmark pairs each image or video item with human-written prompts probing implicit structure that people infer at a glance.
- Items sit under a structured taxonomy covering temporal, spatial, social and abstract reasoning, including whether a vehicle fits between two parked cars.
- The researchers say existing visual benchmarks target either deliberate expert-level analysis or low-level perception, leaving intuitive reasoning largely untested.