Benchmark or evalResearch and knowledge

super-benchmark

Evaluate whether agents can set up and execute tasks from real research repositories.

Stars
53
Forks
4
License
Apache-2.0
Updated
Updated Jun 1, 2026

Checking live repository facts…

Project overview

SUPER evaluates whether LLM agents can configure environments, execute code, and answer questions drawn from real machine-learning and NLP research repositories. Its dataset is divided into Expert, Masked, and AutoGen task sets, with code for running agents and scoring results. Local mode executes repository code directly on the machine, while Docker and Modal backends offer safer isolation and concurrent benchmark runs.

Repository facts

Primary language
Jupyter Notebook
License
Apache-2.0
Repository updated
Jun 1, 2026
Default branch
main

Resource types

Benchmark or eval

Use cases

Research and knowledge

Runtime

Local

Capabilities

Code executionVerification and evals

Related projects

Browse more similar projects