AgentBench
Evaluate LLM agents across diverse interactive environments and tasks.
- Stars
- 3,557
- Forks
- 270
- Updated
- Updated Jul 9, 2026
Evaluate LLM agents across diverse interactive environments and tasks.
AgentBench measures how well LLMs act autonomously across eight interactive environments, including database and operating-system tasks. The current repository carries a function-calling version integrated with AgentRL, while the older benchmark remains available as v0.1. Running the suite requires Docker, and Python 3.9 is recommended because several scientific dependencies are pinned to older versions.
Resource types
Use cases
Runtime
Capabilities
Train and evaluate systems on abstract grid-transformation reasoning tasks.
Evaluate whether agents can set up and execute tasks from real research repositories.
Public GitHub facts last synced Jul 10, 2026.
Run and share interactive Go notebooks in Jupyter or nteract.
Extract websites into clean Markdown, JSON, or LLM-ready context through CLI, MCP, or API.