Project overview
verified-experiments adds a Python harness and Claude skill with guards for data and Git provenance, real-data sanity, label leakage, impossible metrics, stored predictions, reproducibility, silent fallbacks, guard meta-tests, and code review. Each guard rejects a deliberate fake, and the diagnoser explains the failed gate and next step. An optional agent loop may edit experiment code but cannot change the guards, tests, reviewer, or diagnoser to make a run green.
Repository facts
- Primary language
- Python
- License
- MIT
- Repository updated
- Jul 10, 2026
- Default branch
- main
Resource types
General skill
Use cases
Code review and qualityTesting and debugging
Platforms
Claude Code, Codex, and more
Capabilities
Security guardrailVerification and evals