verified-experiments
Guards ML experiments against fake results with provenance, leakage, sanity, and reviewer gates.
Guards ML experiments against fake results with provenance, leakage, sanity, and reviewer gates.
verified-experiments adds a Python harness and Claude skill with guards for data and Git provenance, real-data sanity, label leakage, impossible metrics, stored predictions, reproducibility, silent fallbacks, guard meta-tests, and code review. Each guard rejects a deliberate fake, and the diagnoser explains the failed gate and next step. An optional agent loop may edit experiment code but cannot change the guards, tests, reviewer, or diagnoser to make a run green.
Resource types
Use cases
Platforms
Review high-risk code through independent specialist agents and synthesis.
Audit a repository locally and turn cited findings into clear priorities.
Capabilities
Public GitHub facts last synced Jul 10, 2026.
Review code with adjustable roast levels and actionable fixes.
Score and audit Claude SKILL.md files