vitabench
Evaluates LLM agents on real-world interactive tasks, tool use, and long-horizon application workflows.
- Stars
- 156
- Forks
- 17
- Updated
- Updated Jul 10, 2026
Evaluates LLM agents on real-world interactive tasks, tool use, and long-horizon application workflows.
VitaBench evaluates agents in real-world service scenarios such as delivery, in-store operations, and OTA, using databases, API tools, and 100 tasks per domain. Configure the models, run the vita command, save simulations under data/simulations, and re-evaluate existing runs when needed. The current release updates datasets, tools, evaluator models, and metrics, so scores should be interpreted alongside the chosen configuration.
Resource types
Use cases
Benchmarks autonomous agents on long-horizon real-world tasks with up to million-token contexts.
Run deterministic, cross-model evaluations for agent skills in Docker.
Runtime
Capabilities
Audience
Public GitHub facts last synced Jul 10, 2026.
Evaluates agent trace debugging across reasoning, execution, and planning errors.
Run and score multiple model providers on ARC-AGI tasks.