trail-benchmark
Evaluates agent trace debugging across reasoning, execution, and planning errors.
- Stars
- 20
- Forks
- 1
- Updated
- Updated Jun 27, 2026
Evaluates agent trace debugging across reasoning, execution, and planning errors.
TRAIL is a benchmark dataset with 148 annotated agent execution traces and 841 errors across reasoning, execution, and planning. The traces come from software-engineering and information-retrieval tasks, and the evaluation scripts run a LiteLLM-compatible model on GAIA or SWE Bench splits before calculating scores. It is intended for studying error localization in complex agent workflows.
Resource types
Use cases
Runtime
Evaluates LLM agents on real-world interactive tasks, tool use, and long-horizon application workflows.
Benchmarks autonomous agents on long-horizon real-world tasks with up to million-token contexts.
Capabilities
Audience
Public GitHub facts last synced Jul 10, 2026.
Run deterministic, cross-model evaluations for agent skills in Docker.
Run and score multiple model providers on ARC-AGI tasks.