Project overview
AgentBench measures how well LLMs act autonomously across eight interactive environments, including database and operating-system tasks. The current repository carries a function-calling version integrated with AgentRL, while the older benchmark remains available as v0.1. Running the suite requires Docker, and Python 3.9 is recommended because several scientific dependencies are pinned to older versions.
Repository facts
- Primary language
- Python
- License
- Apache-2.0
- Repository updated
- Jul 9, 2026
- Default branch
- main
Resource types
Benchmark or eval
Use cases
Research and knowledge
Runtime
Docker
Capabilities
Verification and evals