inspect_evals
Run and extend a broad collection of Inspect AI evaluations.
- Stars
- 575
- Forks
- 375
- Updated
- Updated Jul 10, 2026
Run and extend a broad collection of Inspect AI evaluations.
Inspect Evals gathers benchmarks for coding, cybersecurity, safeguards, reasoning, knowledge, and other model capabilities. Researchers can run individual tasks with Inspect AI, improve included evaluations, or register externally maintained implementations. The project recommends Python 3.11 or 3.12 for the most reliable development and execution experience.
Resource types
Use cases
Capabilities
Audience
Generate test automation code across major frameworks and languages.
Reuse practical QA automation agents and testing playbooks across different AI assistants.
Automate and test Chromium, Firefox, and WebKit through one API.
Public GitHub facts last synced Jul 10, 2026.
Automate and test visible interfaces with natural-language instructions.