ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence
The first interactive ARC benchmark: an agent must explore a game with no instructions, scored on action efficiency vs humans — frontier AI scores <1%
Its job in the ARC-AGI-3 storyThe problem definition itself: the first interactive, instruction-free ARC benchmark, where an agent dropped into a novel turn-based game must explore, model dynamics, infer the goal, and plan across four pillars, scored by action efficiency relative to a human baseline (RHAE). It sets the frontier gap the whole shelf exists to close: humans solve 100% while frontier AI scores below 1%.