A The ARC Atlas

Part Ⅳ · The chronological spine

Seven years of ARC-AGI.

ARC-AGI began in 2019 with a new argument about how to measure intelligence. After five years without a decisive machine breakthrough, new methods broke through in 2024—only for ARC-AGI-2 and ARC-AGI-3 to raise the bar again. Every milestone links to its source in Part Ⅱ.

2019–2026
The ARC era
3
Benchmark generations
14
Milestones
2
Resets of the bar

2019 · The measure

A benchmark built to measure the thing scale hides.

Nov 2019

On the Measure of Intelligencepaper

François Chollet defines intelligence as skill-acquisition efficiency, argues that a fair test must control for prior knowledge, and introduces ARC: a few-shot, Core-Knowledge benchmark where every task is new. This becomes the reference point for later work.

№01 · full entry in Part Ⅱ →

2020 · The template

A solo searcher wins, and sets a pattern that holds for three years.

2020

The first Kaggle competition — Icecuber winscompetition

A solo entrant wins with a hand-built DSL of ~142 grid operations, composed through exhaustive search and filtered by an exact-match verifier. With no learning, the DSL + search + verifier template leads for three years while large models stay below 5%.

№07 · full entry in Part Ⅱ →2020 solution write-ups ↗

2024 · The break

Test-time adaptation, then o3 — the record moves at last.

Jun 2024

50% with GPT-4o — sampling as an axismilestone

Ryan Greenblatt samples thousands of GPT-4o programs per task, keeps those that replay every demonstration pair, and votes. The result shows that sample count against a verifier is a tunable axis of accuracy and anticipates the refinement loop.

№10 · full entry in Part Ⅱ →

Nov 2024

Test-time training formalizedpaper

MIT shows that fitting a fresh adapter to each task's demonstrations can produce a roughly 6× gain over a fine-tuned baseline and reach average-human level when combined with program synthesis. Per-task adaptation defines the 2024 season.

№08 · full entry in Part Ⅱ →

2024

The ARChitects win the ARC Prizecompetition

Test-time training on an open 8B model plus a product-of-experts over task symmetries takes the 2024 Kaggle competition — entirely offline, at about two cents per task. The winners report records 53.5% for the prize winner and 55.5% for an ineligible non-open submission.

№09 · full entry in Part Ⅱ →Official 2024 results ↗

Dec 2024

o3 clears 87.5% — the step-functionmilestone

OpenAI's o3, searching over chains-of-thought with a learned evaluator, posts the first human-competitive score on ARC-AGI-1 — at tens to thousands of dollars a task. It is a proprietary benchmark record, not an open-source prize win, and it forces a redesign.

№06 · full entry in Part Ⅱ →Official o3 analysis ↗

2025 · The reset

A harder, cost-scored benchmark, and the first interactive preview.

2025

ARC-AGI-2 reopens the gappaper

A harder benchmark adds compositional tasks with multiple interacting rules and in-context symbols, plus cost per task as a scored axis. Launch-era frontier-model scores fell to the low single digits, reopening the human–machine gap.

№02 · full entry in Part Ⅱ →ARC-AGI-2 report ↗

Jul–Aug 2025

The ARC-AGI-3 Preview Challengecompetition

A 30-day preview of the interactive benchmark. StochasticGoose (12.58%) predicts which actions change the frame; Blind Squirrel (6.71%) builds a state graph. The results establish informed exploration as the first useful baseline.

Official preview learnings ↗

2025

The refinement loop — ARC Prize 2025 reportpaper

The official account of the first ARC-AGI-2 season names the year's theme: propose, verify against the demos, feed back, iterate. NVARC wins at 24.03%; the 85% grand prize goes unclaimed; cheap-and-small beats costly-and-large.

№05 · full entry in Part Ⅱ →Official 2025 results ↗

2025

ARC without pretraining — CompressARCpaper

A different approach solves each task by minimizing description length, with no pretraining or external program search. Its ~20–34% result tests the claim that MDL may underlie what search-and-adapt methods approximate.

№16 · full entry in Part Ⅱ →

2026 · The frontier

The interactive benchmark lands, essentially unsolved.

Dec 2025

Graph-based exploration takes 3rdpaper

The first published ARC-AGI-3 agent is a training-free explorer that keeps a directed graph of hashed frame-states and heads toward the nearest untested action. It uses more of the interaction budget than model-controlled agents, but explores without inducing a transition model.

№17 · full entry in Part Ⅱ →

2026

ARC-AGI-3 — the interactive benchmarkpaper

The first format change since 2019: an agent enters a turn-based game with no instructions or stated goal, using only the controls currently available. RHAE scores action efficiency; humans solve every included environment, while the report's harness-free launch baselines sit below 1%.

№03 · full entry in Part Ⅱ →ARC-AGI-3 technical report ↗

2026

ARC Prize 2026 — $2M on Kagglecompetition

Two tracks run on Kaggle — ARC-AGI-3 and a final ARC-AGI-2 track — with a $2M pool. The ARC-AGI-3 season publishes two paid milestones before the final awards.

Official 2026 schedule ↗

Jul 2026

The Duck wins ARC-AGI-3 Milestone #1milestone

Tufa Labs' small local coding agent takes first, with Reki second and forge third. The podium shifts from preview-era graph search toward vision-language policies with memory, structured actions, and—in The Duck—a live Python REPL. The final competition remains open.

Official milestone report ↗

The shape of it

Every time the machines catch up, the bar moves.

Across the three competitions, the pattern repeats: a method closes the gap, then a redesigned benchmark opens a harder one. Icecuber's searchSolo, Icecuber's hand-built DSL of ~142 grid operations — greedily composed into a DAG, no learning at all — cleared ~20% of the private set; pooling every 2020 entry into one ensemble reached only ~49%. That was the ceiling of pure search. yielded to test-time trainingMIT's recipe fits a fresh LoRA adapter to each task's own demonstrations at inference — up to a 6× accuracy lever, and ~62% once ensembled with program synthesis, matching average-human performance., then o3o3 scored 75.7% at high efficiency and 87.5% using ~172× the compute — at roughly $20 to thousands of dollars a task, against ~$5 for a human. Chollet's own verdict: not AGI.. o3 prompted ARC-AGI-2; ARC-AGI-2 led to the interactive ARC-AGI-3.

Figure · The climb and the resets

The frontier score across all three generations

Each reset drops the best-known score back toward zero against a near-100% human baseline.

The climb and resets Selected frontier scores across the three ARC-AGI benchmark generations. 0 25 50 75 100'20'21'22'23'24'25'26 human ≈ 98% 20% 42% 55.5% 87.5%ARC-AGI-1 3% 24.03%ARC-AGI-2 0.5% 12.58%ARC-AGI-3
ARC-AGI-1ARC-AGI-2ARC-AGI-3 Human ≈ 98–100%

Five eras, one idea

2020

Hand-built search

Icecuber

A solo DSL + brute-force search wins the first Kaggle. No learning; the template holds for 3 years.

2024

Test-time adaptation

MindsAI · MIT · ARChitects

Per-task fine-tuning and augmentation-ensembling drive the private record from 33% to 55.5%.

Dec 2024

The o3 break

OpenAI

Guided search over chains-of-thought clears 87.5% — near-human, but at thousands of dollars a task.

2025

The refinement loop

ARC-AGI-2 field

Propose → verify → feed back → iterate. Cheap small models beat costly big ones. Top score: 24%.

2025+

Agentic exploration

ARC-AGI-3 preview + 2026 milestones

Interactive worlds. Exploration progresses from state graphs to multimodal policies, memory, and executable tools; transferable world modeling and goal induction remain open.

The techniques change, but the recurring pattern is adaptation followed by verification. ARC-AGI-3 removes the clean verifier: there are no demonstration pairs or stated goals, only a world to probe.

→ Who won each round, in Part Ⅸ · Winning Solutions

YouTube first · local synths follow

ARC Radio

01 / 12 🦉 8-Bit Chiptune Playlist 🦉 Retro Video Game Music for Nostalgic Vibes YouTube · external stream

The 4 requested YouTube selections play first and require a network connection; their titles refresh from YouTube when they load. 8 original AI-composed retro-game loops follow and are generated live in your browser. Audio keeps playing when you close this panel and stops only when you press Pause.

Field notes · reader review

Help improve this guide

Found a wrong score, broken link, missing paper, or unclear passage? Tell us what you noticed.

How useful is it? optional
- / 5
What kind of note? optional

No account, no tracking. Sent straight to the maintainer.