agent-skills-eval
Compare skill-enabled and baseline runs to measure whether an Agent Skill actually helps.
Compare skill-enabled and baseline runs to measure whether an Agent Skill actually helps.
agent-skills-eval tests an Agent Skill against the same prompts in two modes: with the skill loaded and without it as a baseline. A judge model grades both outputs against defined assertions, producing JSON and JSONL artifacts plus a static HTML report that shows measurable lift, failures, timing, and tool-call behavior. The project includes a command-line runner and a TypeScript SDK, supports OpenAI-compatible chat backends and custom providers, and can run in CI. Users supply both the target model and the judge model.
Resource types
Use cases
Runtime
Automate and test Chromium, Firefox, and WebKit through one API.
Control, automate, and operate Android devices or fleets from one integrated platform.
Capabilities
Public GitHub facts last synced Jul 14, 2026.
Configure and run web, mobile, and load tests through one QA command line.
Choose an architecture before creating, evaluating, improving, and packaging reusable AI skills.