Project overview
agent-skills-eval tests an Agent Skill against the same prompts in two modes: with the skill loaded and without it as a baseline. A judge model grades both outputs against defined assertions, producing JSON and JSONL artifacts plus a static HTML report that shows measurable lift, failures, timing, and tool-call behavior. The project includes a command-line runner and a TypeScript SDK, supports OpenAI-compatible chat backends and custom providers, and can run in CI. Users supply both the target model and the judge model.
Repository facts
- Primary language
- TypeScript
- License
- MIT
- Repository updated
- Jul 13, 2026
- Default branch
- main
Resource types
FrameworkCLI appSDK
Use cases
Testing and debuggingAI and agent development
Runtime
Command lineLocal
Capabilities
Verification and evals