ARC-AGI-2
Train and evaluate systems on abstract grid-transformation reasoning tasks.
- Stars
- 723
- Forks
- 106
- Updated
- Updated Jul 10, 2026
Train and evaluate systems on abstract grid-transformation reasoning tasks.
ARC-AGI-2 contains 1,000 public training tasks and 120 public evaluation tasks built from small colored-integer grids. Solvers infer a transformation from demonstration pairs and must produce every test output with the correct dimensions and cells. The evaluation set is intended for previously unseen testing, so repeatedly tuning against it would compromise the result.
Resource types
Use cases
Capabilities
Evaluate LLM agents across diverse interactive environments and tasks.
Evaluate whether agents can set up and execute tasks from real research repositories.
Public GitHub facts last synced Jul 10, 2026.
Run and share interactive Go notebooks in Jupyter or nteract.
Extract websites into clean Markdown, JSON, or LLM-ready context through CLI, MCP, or API.