Project overview
AgencyBench V2 covers six capabilities across game, frontend, backend, code, research, and MCP work, with 32 real-world long-horizon scenarios and 138 tasks. A user-simulation agent supplies iterative feedback while a Docker sandbox runs functional and visual checks using rule-based, vision-based, and LLM judges. Scenarios require model and evaluator configuration, and game or frontend cases also need the remote Docker sandbox.
Repository facts
- Primary language
- Python
- License
- MIT
- Repository updated
- Jul 10, 2026
- Default branch
- main
Resource types
Benchmark or evalDatasetFramework
Use cases
Testing and debuggingAI and agent developmentResearch and knowledge
Runtime
DockerLocalSandboxed
Protocols & integrations
Model Context Protocol
Capabilities
Computer useData visualizationHuman in the loopVerification and evals
Audience
DevelopersResearchers