Project overview
This benchmarking toolkit runs ARC-AGI tasks through configurable adapters for providers such as OpenAI, Anthropic, and Gemini, with rate limiting, retries, saved submissions, and scoring built in. It supports both single-task debugging and asynchronous batch runs. Full evaluations require separately cloned ARC-AGI data, and hosted models require their respective API credentials.
Repository facts
- Primary language
- Python
- License
- MIT
- Repository updated
- Jul 3, 2026
- Default branch
- main
Resource types
Benchmark or eval
Use cases
Testing and debugging
Runtime
Command lineLocal