inspect_ai
Design LLM evaluations with tool use, multi-turn dialogue, and model grading.
- Stars
- 2,328
- Forks
- 595
- Updated
- Updated Jul 10, 2026
Design LLM evaluations with tool use, multi-turn dialogue, and model grading.
Inspect is an evaluation framework created by the UK AI Security Institute for testing large language models. It includes components for prompt construction, tool use, multi-turn dialogue, and model-graded evaluation, plus more than 200 pre-built evaluations. Researchers can compose these capabilities into tasks and extend the framework through Python packages that add new elicitation or scoring techniques.
Resource types
Use cases
Capabilities
Public GitHub facts last synced Jul 10, 2026.
Build agents, workflows, and deployable AI applications with TypeScript.
Build extensible browser and desktop IDEs for multiple languages.
Build and ship compact cross-platform desktop apps with TypeScript.
Build and package cross-platform desktop apps with Next.js and Electron.