AgencyBench
Benchmarks autonomous agents on long-horizon real-world tasks with up to million-token contexts.
- Stars
- 89
- Forks
- 4
- Updated
- Updated Jul 10, 2026
Benchmarks autonomous agents on long-horizon real-world tasks with up to million-token contexts.
AgencyBench V2 covers six capabilities across game, frontend, backend, code, research, and MCP work, with 32 real-world long-horizon scenarios and 138 tasks. A user-simulation agent supplies iterative feedback while a Docker sandbox runs functional and visual checks using rule-based, vision-based, and LLM judges. Scenarios require model and evaluator configuration, and game or frontend cases also need the remote Docker sandbox.
Resource types
Use cases
Evaluates LLM agents on real-world interactive tasks, tool use, and long-horizon application workflows.
Run deterministic, cross-model evaluations for agent skills in Docker.
Runtime
Protocols & integrations
Capabilities
Audience
Public GitHub facts last synced Jul 10, 2026.
Evaluates agent trace debugging across reasoning, execution, and planning errors.
Run and score multiple model providers on ARC-AGI tasks.