waza
Create skill evaluations, run benchmarks, and compare performance across models.
- Stars
- 1,088
- Forks
- 65
- Updated
- Updated Jul 14, 2026
Create skill evaluations, run benchmarks, and compare performance across models.
Waza is a Go CLI for scaffolding Agent Skills evaluation suites, running benchmarks, and comparing outcomes across models or executors. Its readiness checks cover frontmatter compliance, token budgets, evaluation files, and agentskills.io specification rules, and the workflow can be integrated into CI. It is aimed at skill authors and teams that need repeatable quality measurements and regression checks.
Resource types
Use cases
Platforms
Run repeatable evaluations of whether agents discover and use a skill correctly.
Traces AI-agent processes, files, network destinations, and model calls with eBPF.
Runtime
Capabilities
Public GitHub facts last synced Jul 14, 2026.
Fetch, merge, and filter Databricks job logs from one CLI command.
Diagnose deployed web apps behind enterprise authentication from the terminal.