General skillAI and agent development

llm-judge

Design runnable LLM evaluations with rubrics, judge prompts, and Python harnesses.

Stars
0
Forks
0
License
MIT
Updated
Updated Apr 23, 2026

Checking live repository facts…

Project overview

The skill moves from scope to measurable rubrics, bias-aware judge prompts, and a runnable Python harness. Generated tooling supports content-hash caching, watch mode, per-step trajectory attribution, and calibration against human labels. The standalone example uses an Anthropic API key, while Claude Code guides the design and scaffolding.

Repository facts

Primary language
Python
License
MIT
Repository updated
Apr 23, 2026
Default branch
main

Resource types

General skill

Use cases

AI and agent development

Platforms

Claude Code, Codex, and more

Capabilities

Verification and evals

Related projects

Browse more similar projects