Benchmark or evalTesting and debugging

skill-optimizer

Run deterministic, cross-model evaluations for agent skills in Docker.

Stars
70
Forks
11
License
MIT
Updated
Updated Jul 7, 2026

Checking live repository facts…

Project overview

skill-optimizer provides both an agent skill for authoring eval suites and a local CLI for executing them in Docker against OpenRouter models. Cases and suites can be repeated across trials to benchmark skill behavior, debug failures, and compare reliability across models. The workbench can also expose hidden services such as MCP servers to the agent during a controlled test.

Repository facts

Primary language
TypeScript
License
MIT
Repository updated
Jul 7, 2026
Default branch
development

Resource types

Benchmark or evalPlugin

Use cases

Testing and debugging

Platforms

Claude CodeCursorCodexOpenCode

Runtime

Command lineDockerLocal

Capabilities

Verification and evals

Related projects

Browse more similar projects