Benchmark or evalResearch and knowledge

AgentBench

Evaluate LLM agents across diverse interactive environments and tasks.

Stars
3,557
Forks
270
License
Apache-2.0
Updated
Updated Jul 9, 2026

Checking live repository facts…

Project overview

AgentBench measures how well LLMs act autonomously across eight interactive environments, including database and operating-system tasks. The current repository carries a function-calling version integrated with AgentRL, while the older benchmark remains available as v0.1. Running the suite requires Docker, and Python 3.9 is recommended because several scientific dependencies are pinned to older versions.

Repository facts

Primary language
Python
License
Apache-2.0
Repository updated
Jul 9, 2026
Default branch
main

Resource types

Benchmark or eval

Use cases

Research and knowledge

Runtime

Docker

Capabilities

Verification and evals

Related projects

Browse more similar projects