Benchmark or evalTesting and debugging

AgencyBench

Benchmarks autonomous agents on long-horizon real-world tasks with up to million-token contexts.

Stars
89
Forks
4
License
MIT
Updated
Updated Jul 10, 2026

Checking live repository facts…

Project overview

AgencyBench V2 covers six capabilities across game, frontend, backend, code, research, and MCP work, with 32 real-world long-horizon scenarios and 138 tasks. A user-simulation agent supplies iterative feedback while a Docker sandbox runs functional and visual checks using rule-based, vision-based, and LLM judges. Scenarios require model and evaluator configuration, and game or frontend cases also need the remote Docker sandbox.

Repository facts

Primary language
Python
License
MIT
Repository updated
Jul 10, 2026
Default branch
main

Resource types

Benchmark or evalDatasetFramework

Use cases

Testing and debuggingAI and agent developmentResearch and knowledge

Runtime

DockerLocalSandboxed

Protocols & integrations

Model Context Protocol

Capabilities

Computer useData visualizationHuman in the loopVerification and evals

Audience

DevelopersResearchers

Related projects

Browse more similar projects