llm-judge
Design runnable LLM evaluations with rubrics, judge prompts, and Python harnesses.
- Stars
- 0
- Forks
- 0
- Updated
- Updated Apr 23, 2026
Design runnable LLM evaluations with rubrics, judge prompts, and Python harnesses.
The skill moves from scope to measurable rubrics, bias-aware judge prompts, and a runnable Python harness. Generated tooling supports content-hash caching, watch mode, per-step trajectory attribution, and calibration against human labels. The standalone example uses an Anthropic API key, while Claude Code guides the design and scaffolding.
Resource types
Use cases
Platforms
Capabilities
Turn detailed business processes into reusable, tested AI skills.
Define portable task-specific sub-agents in Markdown for several coding assistants.
Public GitHub facts last synced Jul 10, 2026.
Find, create, run, and improve agent skills from real execution feedback.
Reduce LLM token use by pruning context, caching prompts, and routing models.