skillgrade
Run repeatable evaluations of whether agents discover and use a skill correctly.
- Stars
- 618
- Forks
- 45
- Updated
- Updated Jul 14, 2026
Run repeatable evaluations of whether agents discover and use a skill correctly.
Skillgrade scaffolds an eval.yaml for a SKILL.md directory, runs repeated tasks, and measures whether an agent discovers and applies the skill as intended. It supports Gemini, Claude, Codex, ACP, OpenCode, and custom command agents, with deterministic or LLM-rubric graders and CLI or browser reports. Node.js 20+ is required, while Docker, API credentials, or a local command depend on the selected executor.
Resource types
Use cases
Platforms
Create skill evaluations, run benchmarks, and compare performance across models.
Traces AI-agent processes, files, network destinations, and model calls with eBPF.
Runtime
Protocols & integrations
Capabilities
Public GitHub facts last synced Jul 14, 2026.
Fetch, merge, and filter Databricks job logs from one CLI command.
Diagnose deployed web apps behind enterprise authentication from the terminal.