eval-harness-architect
Builds regression-ready eval harnesses for LLM, agent, and RAG systems.
Builds regression-ready eval harnesses for LLM, agent, and RAG systems.
eval-harness-architect starts from the decision an evaluation must support, then builds versioned datasets, holdouts, dimension-specific scorers, and a calibrated judge. It measures significance, wires regression gates and PR scorecards into CI, and adds drift and production-feedback loops without fabricating labels or benchmark numbers.
Resource types
Use cases
Platforms
Capabilities
Test mobile and desktop app navigation without routine screenshots.
Capture and decode logic-analyzer signals through natural-language requests.
Public GitHub facts last synced Jul 10, 2026.
Turn one browser exploration into repeatable end-to-end tests.
Query CloudWatch in natural language and investigate AWS incidents from Claude Code.