Project overview
skill-optimizer provides both an agent skill for authoring eval suites and a local CLI for executing them in Docker against OpenRouter models. Cases and suites can be repeated across trials to benchmark skill behavior, debug failures, and compare reliability across models. The workbench can also expose hidden services such as MCP servers to the agent during a controlled test.
Repository facts
- Primary language
- TypeScript
- License
- MIT
- Repository updated
- Jul 7, 2026
- Default branch
- development
Resource types
Benchmark or evalPlugin
Use cases
Testing and debugging
Platforms
Claude CodeCursorCodexOpenCode
Runtime
Command lineDockerLocal
Capabilities
Verification and evals