Benchmark or evalTesting and debugging

trail-benchmark

Evaluates agent trace debugging across reasoning, execution, and planning errors.

Stars
20
Forks
1
License
MIT
Updated
Updated Jun 27, 2026

Checking live repository facts…

Project overview

TRAIL is a benchmark dataset with 148 annotated agent execution traces and 841 errors across reasoning, execution, and planning. The traces come from software-engineering and information-retrieval tasks, and the evaluation scripts run a LiteLLM-compatible model on GAIA or SWE Bench splits before calculating scores. It is intended for studying error localization in complex agent workflows.

Repository facts

Primary language
Python
License
MIT
Repository updated
Jun 27, 2026
Default branch
main

Resource types

Benchmark or evalDataset

Use cases

Testing and debuggingAI and agent development

Runtime

Command lineLocal

Capabilities

Verification and evals

Audience

DevelopersResearchers

Related projects

Browse more similar projects