skill-optimizer
Run deterministic, cross-model evaluations for agent skills in Docker.
- Stars
- 70
- Forks
- 11
- Updated
- Updated Jul 7, 2026
Run deterministic, cross-model evaluations for agent skills in Docker.
skill-optimizer provides both an agent skill for authoring eval suites and a local CLI for executing them in Docker against OpenRouter models. Cases and suites can be repeated across trials to benchmark skill behavior, debug failures, and compare reliability across models. The workbench can also expose hidden services such as MCP servers to the agent during a controlled test.
Resource types
Use cases
Platforms
Evaluates LLM agents on real-world interactive tasks, tool use, and long-horizon application workflows.
Benchmarks autonomous agents on long-horizon real-world tasks with up to million-token contexts.
Runtime
Capabilities
Public GitHub facts last synced Jul 10, 2026.
Evaluates agent trace debugging across reasoning, execution, and planning errors.
Run and score multiple model providers on ARC-AGI tasks.