Benchmark or evalTesting and debugging

vitabench

Evaluates LLM agents on real-world interactive tasks, tool use, and long-horizon application workflows.

Stars
156
Forks
17
License
MIT
Updated
Updated Jul 10, 2026

Checking live repository facts…

Project overview

VitaBench evaluates agents in real-world service scenarios such as delivery, in-store operations, and OTA, using databases, API tools, and 100 tasks per domain. Configure the models, run the vita command, save simulations under data/simulations, and re-evaluate existing runs when needed. The current release updates datasets, tools, evaluator models, and metrics, so scores should be interpreted alongside the chosen configuration.

Repository facts

Primary language
Python
License
MIT
Repository updated
Jul 10, 2026
Default branch
main

Resource types

Benchmark or evalCLI appDatasetFramework

Use cases

Testing and debuggingAI and agent developmentPersonal and daily lifeResearch and knowledge

Runtime

Command lineLocal

Capabilities

Data visualizationHuman in the loopTool useVerification and evalsWorkflow automation

Audience

DevelopersResearchers

Related projects

Browse more similar projects