Project overview
Inspect Evals gathers benchmarks for coding, cybersecurity, safeguards, reasoning, knowledge, and other model capabilities. Researchers can run individual tasks with Inspect AI, improve included evaluations, or register externally maintained implementations. The project recommends Python 3.11 or 3.12 for the most reliable development and execution experience.
Repository facts
- Primary language
- Python
- License
- MIT
- Repository updated
- Jul 10, 2026
- Default branch
- main
Resource types
Collection or directory
Use cases
Testing and debugging
Capabilities
Verification and evals
Audience
Developers