EvoSkill
An open-source framework that evolves coding-agent skills from failed benchmark trajectories.
- Stars
- 1,035
- Forks
- 111
- Updated
- Updated Jul 14, 2026
An open-source framework that evolves coding-agent skills from failed benchmark trajectories.
EvoSkill is an open-source framework for discovering and evolving reusable coding-agent skills. It runs agents on benchmark tasks, analyzes failed trajectories, proposes skill and system-prompt variants, and evaluates them on held-out data. Projects can use Claude Code, Codex, OpenCode, OpenHands, or Goose with CSV question-and-answer sets or containerized Harbor tasks. It suits teams with a meaningful evaluation set and requires the selected agent runtime, model-provider credentials, and supporting environment.
Resource types
Use cases
Platforms
Combine interchangeable models, data sources, and tools in LLM applications.
Trace, evaluate, monitor, and manage LLM, agent, and ML workflows.
Runtime
Public GitHub facts last synced Jul 14, 2026.
Build and run Python agents with graph-based workflows
Build type-safe Python agents with tools, structured outputs, and provider choice