FrameworkTesting and debugging

agent-skills-eval

Compare skill-enabled and baseline runs to measure whether an Agent Skill actually helps.

Stars
620
Forks
33
License
MIT
Updated
Updated Jul 13, 2026

Checking live repository facts…

Project overview

agent-skills-eval tests an Agent Skill against the same prompts in two modes: with the skill loaded and without it as a baseline. A judge model grades both outputs against defined assertions, producing JSON and JSONL artifacts plus a static HTML report that shows measurable lift, failures, timing, and tool-call behavior. The project includes a command-line runner and a TypeScript SDK, supports OpenAI-compatible chat backends and custom providers, and can run in CI. Users supply both the target model and the judge model.

Repository facts

Primary language
TypeScript
License
MIT
Repository updated
Jul 13, 2026
Default branch
main

Resource types

FrameworkCLI appSDK

Use cases

Testing and debuggingAI and agent development

Runtime

Command lineLocal

Capabilities

Verification and evals

Related projects

Browse more similar projects