Benchmark or evalTesting and debugging

arc-agi-benchmarking

Run and score multiple model providers on ARC-AGI tasks.

Stars
351
Forks
69
License
MIT
Updated
Updated Jul 3, 2026

Checking live repository facts…

Project overview

This benchmarking toolkit runs ARC-AGI tasks through configurable adapters for providers such as OpenAI, Anthropic, and Gemini, with rate limiting, retries, saved submissions, and scoring built in. It supports both single-task debugging and asynchronous batch runs. Full evaluations require separately cloned ARC-AGI data, and hosted models require their respective API credentials.

Repository facts

Primary language
Python
License
MIT
Repository updated
Jul 3, 2026
Default branch
main

Resource types

Benchmark or eval

Use cases

Testing and debugging

Runtime

Command lineLocal

Related projects

Browse more similar projects