Frontier snapshot · August 2026

The human line

Each dot is a real task. The diagonal is parity with people. Points above it are where machines already beat the human baseline. Points below it are where people still lead — often by a lot.

Machines ahead
15
Roughly even
3
People ahead
12

AI score

Human score

Reasoning

Novel interactive games (ARC-AGI-3)

People still ahead

Human

100

Untrained people on calibrated games

Frontier AI

30

Claude Opus 5 (most models <8%)

Gap -70 points on a 0–100 task scale

Give a person a new game with no manual and they learn it. Give a frontier model the same world and it mostly flails. This is the cleanest remaining gap in fluid skill-acquisition.

Interactive, no language crutches, scored against human action efficiency. The public board is still sparse.

Try an item from this test

Sit the test

One item from every benchmark on the map. Puzzles are grids and boards. Official items are marked; the rest are faithful to how the real test is taken, because the held-out sets cannot be published.

0 of 30 tried

Knowledge

Math

Code

Reasoning

Games

Perception

Science

Language & art

Physical world

People still aheadSame format as the testARC-AGI-3 format

A game with no manual

ARC-AGI-3 scores how efficiently a player acquires a new game. People do this. Models mostly flail.

You are the pale square. Reach the marked cell. Click an adjacent tile to step. That is all you are told — same as the real test.

Adjacent tiles only.

Every task in this view

Tap a row to pin it on the chart.