super-benchmark
Evaluate whether agents can set up and execute tasks from real research repositories.
- Stars
- 53
- Forks
- 4
- Updated
- Updated Jun 1, 2026
Evaluate whether agents can set up and execute tasks from real research repositories.
SUPER evaluates whether LLM agents can configure environments, execute code, and answer questions drawn from real machine-learning and NLP research repositories. Its dataset is divided into Expert, Masked, and AutoGen task sets, with code for running agents and scoring results. Local mode executes repository code directly on the machine, while Docker and Modal backends offer safer isolation and concurrent benchmark runs.
Resource types
Use cases
Runtime
Capabilities
Evaluate LLM agents across diverse interactive environments and tasks.
Train and evaluate systems on abstract grid-transformation reasoning tasks.
Public GitHub facts last synced Jul 10, 2026.
Run and share interactive Go notebooks in Jupyter or nteract.
Extract websites into clean Markdown, JSON, or LLM-ready context through CLI, MCP, or API.