Project overview
SUPER evaluates whether LLM agents can configure environments, execute code, and answer questions drawn from real machine-learning and NLP research repositories. Its dataset is divided into Expert, Masked, and AutoGen task sets, with code for running agents and scoring results. Local mode executes repository code directly on the machine, while Docker and Modal backends offer safer isolation and concurrent benchmark runs.
Repository facts
- Primary language
- Jupyter Notebook
- License
- Apache-2.0
- Repository updated
- Jun 1, 2026
- Default branch
- main
Resource types
Benchmark or eval
Use cases
Research and knowledge
Runtime
Local
Capabilities
Code executionVerification and evals