Collection or directoryTesting and debugging

inspect_evals

Run and extend a broad collection of Inspect AI evaluations.

Stars
575
Forks
375
License
MIT
Updated
Updated Jul 10, 2026

Checking live repository facts…

Project overview

Inspect Evals gathers benchmarks for coding, cybersecurity, safeguards, reasoning, knowledge, and other model capabilities. Researchers can run individual tasks with Inspect AI, improve included evaluations, or register externally maintained implementations. The project recommends Python 3.11 or 3.12 for the most reliable development and execution experience.

Repository facts

Primary language
Python
License
MIT
Repository updated
Jul 10, 2026
Default branch
main

Resource types

Collection or directory

Use cases

Testing and debugging

Capabilities

Verification and evals

Audience

Developers

Related projects

Browse more similar projects