deepeval-bcg
Evaluate strategic AI outputs with consulting-style scoring and adversarial checks.
- Stars
- 1
- Forks
- 0
- Updated
- Updated May 21, 2026
Evaluate strategic AI outputs with consulting-style scoring and adversarial checks.
Run the skill against a strategic deliverable to receive a PASS, REVISE, or FAIL verdict, scores across eight consulting-oriented dimensions, and a concrete repair directive. Its Skeptic Agent probes for silent ambiguity choices, agreement with flawed premises, and missing counterarguments, while a separate novelty stack checks whether the insight is more than generic strategy language. The evaluation runs inside Claude Code without separate provider API keys.
Resource types
Use cases
Platforms
Capabilities
Turn detailed business processes into reusable, tested AI skills.
Define portable task-specific sub-agents in Markdown for several coding assistants.
Public GitHub facts last synced Jul 10, 2026.
Find, create, run, and improve agent skills from real execution feedback.
Reduce LLM token use by pruning context, caching prompts, and routing models.