FrameworkSoftware development

inspect_ai

Design LLM evaluations with tool use, multi-turn dialogue, and model grading.

Stars
2,328
Forks
595
License
MIT
Updated
Updated Jul 10, 2026

Checking live repository facts…

Project overview

Inspect is an evaluation framework created by the UK AI Security Institute for testing large language models. It includes components for prompt construction, tool use, multi-turn dialogue, and model-graded evaluation, plus more than 200 pre-built evaluations. Researchers can compose these capabilities into tasks and extend the framework through Python packages that add new elicitation or scoring techniques.

Repository facts

Primary language
Python
License
MIT
Repository updated
Jul 10, 2026
Default branch
main

Resource types

Framework

Use cases

Software development

Capabilities

Verification and evals

Related projects

Browse more similar projects