Project overview
llm-cost-guard statically reviews LLM-calling code for repeated uncached prefixes, oversized model tiers, bloated context, unbounded output, unbudgeted agent loops, and repeated embeddings. It estimates the dominant tokens-times-price-times-volume term, then proposes prompt caching, model routing, context trimming, output caps, or batching while naming latency and quality trade-offs. It has no billing access, and projected savings still require current prices and evaluation results.
Repository facts
- Primary language
- Not detected
- License
- MIT
- Repository updated
- Jul 10, 2026
- Default branch
- main
Resource types
General skill
Use cases
Code review and quality
Platforms
Claude Code, Codex, and more