arc-agi-benchmarking
Run and score multiple model providers on ARC-AGI tasks.
- Stars
- 351
- Forks
- 69
- Updated
- Updated Jul 3, 2026
Run and score multiple model providers on ARC-AGI tasks.
This benchmarking toolkit runs ARC-AGI tasks through configurable adapters for providers such as OpenAI, Anthropic, and Gemini, with rate limiting, retries, saved submissions, and scoring built in. It supports both single-task debugging and asynchronous batch runs. Full evaluations require separately cloned ARC-AGI data, and hosted models require their respective API credentials.
Resource types
Use cases
Runtime
Public GitHub facts last synced Jul 10, 2026.
Evaluates LLM agents on real-world interactive tasks, tool use, and long-horizon application workflows.
Benchmarks autonomous agents on long-horizon real-world tasks with up to million-token contexts.
Run deterministic, cross-model evaluations for agent skills in Docker.
Evaluates agent trace debugging across reasoning, execution, and planning errors.