CLI appTesting and debugging

skillgrade

Run repeatable evaluations of whether agents discover and use a skill correctly.

Stars
618
Forks
45
License
MIT
Updated
Updated Jul 14, 2026

Checking live repository facts…

Project overview

Skillgrade scaffolds an eval.yaml for a SKILL.md directory, runs repeated tasks, and measures whether an agent discovers and applies the skill as intended. It supports Gemini, Claude, Codex, ACP, OpenCode, and custom command agents, with deterministic or LLM-rubric graders and CLI or browser reports. Node.js 20+ is required, while Docker, API credentials, or a local command depend on the selected executor.

Repository facts

Primary language
TypeScript
License
MIT
Repository updated
Jul 14, 2026
Default branch
main

Resource types

CLI appBenchmark or eval

Use cases

Testing and debugging

Platforms

CodexOpenCode

Runtime

Command lineDockerLocal

Protocols & integrations

Agent Client Protocol

Capabilities

Verification and evals

Related projects

Browse more similar projects