Project overview
eval-harness-architect starts from the decision an evaluation must support, then builds versioned datasets, holdouts, dimension-specific scorers, and a calibrated judge. It measures significance, wires regression gates and PR scorecards into CI, and adds drift and production-feedback loops without fabricating labels or benchmark numbers.
Repository facts
- Primary language
- Python
- License
- MIT
- Repository updated
- Jun 23, 2026
- Default branch
- main
Resource types
General skill
Use cases
Testing and debugging
Platforms
Claude Code, Codex, and more
Capabilities
ObservabilityVerification and evals