Reasoning
Novel interactive games (ARC-AGI-3)
Human
100
Untrained people on calibrated games
Frontier AI
30
Claude Opus 5 (most models <8%)
Gap -70 points on a 0–100 task scale
Give a person a new game with no manual and they learn it. Give a frontier model the same world and it mostly flails. This is the cleanest remaining gap in fluid skill-acquisition.
Interactive, no language crutches, scored against human action efficiency. The public board is still sparse.
Try an item from this test