AgencyBench
Benchmarks autonomous agents on long-horizon real-world tasks with up to million-token contexts.
- Stars
- 89
- Forks
- 4
- License
- MIT
Start with a real need, then filter by use case, format, runtime, and repository facts to find projects you can use, combine, or build on.
Benchmarks autonomous agents on long-horizon real-world tasks with up to million-token contexts.
Evaluates LLM agents on real-world interactive tasks, tool use, and long-horizon application workflows.
Train and evaluate systems on abstract grid-transformation reasoning tasks.
Train and evaluate software-engineering agents on real Python repository tasks.
Evaluates agent trace debugging across reasoning, execution, and planning errors.
Compose reusable pattern recipes for seamless AI wallpaper tiles.