Project overview
TRAIL is a benchmark dataset with 148 annotated agent execution traces and 841 errors across reasoning, execution, and planning. The traces come from software-engineering and information-retrieval tasks, and the evaluation scripts run a LiteLLM-compatible model on GAIA or SWE Bench splits before calculating scores. It is intended for studying error localization in complex agent workflows.
Repository facts
- Primary language
- Python
- License
- MIT
- Repository updated
- Jun 27, 2026
- Default branch
- main
Resource types
Benchmark or evalDataset
Use cases
Testing and debuggingAI and agent development
Runtime
Command lineLocal
Capabilities
Verification and evals
Audience
DevelopersResearchers