Open vitabenchBenchmark or evalTesting and debugging
vitabench↗
@meituan-longcat·Python
Evaluates LLM agents on real-world interactive tasks, tool use, and long-horizon application workflows.
Project overviewVitaBench evaluates agents in real-world service scenarios such as delivery, in-store operations, and OTA, using databases, API tools, and 100 tasks per domain. Configure the models, run the vita command, save simulations under data/simulations, and re-evaluate existing runs when needed. The current release updates datasets, tools, evaluator models, and metrics, so scores should be interpreted alongside the chosen configuration.
- Stars
- ★ 156
- Forks
- ⑂ 17
- License
- MIT
Open fable-workflow-skillGeneral skillAI and agent development
fable-workflow-skill↗
@joey114132·Python
Surface unknowns and assumptions before an AI agent starts building.
Project overviewFable Workflow teaches coding agents to inspect the real codebase and expose unspecified decisions before implementation instead of guessing silently. During the work it records assumptions and deviations, then asks for verification evidence before declaring completion. The Claude Code plugin adds advisory hooks and an optional strict stop gate, though its heuristics can occasionally misfire.
- Stars
- ★ 2
- Forks
- ⑂ 0
- License
- MIT
Open imagegen-skillsGeneral skillDesign and creative work
imagegen-skills↗
@veryCoolTimo·Python
Expand a one-line visual idea into a model-ready image prompt.
Project overviewimage-prompt takes a brief idea, identifies the visual archetype, and expands it into a structured prompt covering composition zones, typography, palette, mood, and finish for posters, ads, UI mockups, game art, logos, and more. It can adapt the same concept for several image models and reuse private style presets; image generation is optional, requires a configured provider key, and may incur API charges.
- Stars
- ★ 3
- Forks
- ⑂ 0
- License
- MIT
Open dvalincodeCoding agentSoftware development
dvalincode↗
@arthurpanhku·TypeScript
Run a local-first, policy-bound coding agent with approvals and tamper-evident audits.
Project overviewDvalinCode targets regulated teams that need AI coding under explicit policy, local data controls, and reviewable run records. Chat stays read-only, Cowork previews and asks for each write, while Code can run routines; the runtime also imports scanner findings into isolated remediation worktrees and supports OpenAI-compatible or local models.
- Stars
- ★ 86
- Forks
- ⑂ 8
- License
- MIT
Open kodoFrameworkSoftware developmentClaude Code
kodo↗
@ikamensh·Python
Orchestrate multiple coding agents for unattended implementation, testing, and review.
Project overviewKodo assigns a goal to coding backends such as Claude Code, Cursor, Codex, Gemini, or Kimi and coordinates repeated work cycles with independent verification. It can implement features, run improvement reviews, test software through realistic user workflows, and fix findings from earlier runs, with effort levels controlling how aggressively agents iterate. At least one supported coding-agent backend must be installed.
- Stars
- ★ 116
- Forks
- ⑂ 6
- License
- MIT
Open tower-defense-skillWeb appPersonal and daily life
tower-defense-skill↗
@mars-tw·JavaScript
Plays a dependency-free Canvas tower defense game with elemental towers, heroes, and boss waves.
Project overviewtower-defense-skill is a native Canvas 2D tower-defense game built around protecting a goddess core. Choose from two maps and three difficulties, build and upgrade six tower types, exploit fire, ice, and lightning matchups, and survive enemy and boss waves. Souls earned during waves unlock hero draws, while deployed heroes can receive a patrol radius; results include progression, achievements, and map-specific rankings.
- Stars
- ★ 2
- Forks
- ⑂ 1
- License
- MIT
Open fable-token-saverGeneral skillSoftware development
fable-token-saver↗
@vincemakes·Shell
Delegate large coding tasks while reserving top-tier models for judgment.
Project overviewUse this Claude Code skill for large, well-specified implementation work that can be divided into task packets. The main model handles decomposition and review, cheaper workers write code, and type checks or tests gate their output. Lite mode reduces top-tier quota use at relatively low total cost; max mode saves more top-tier quota but raised total dollar cost in the benchmark. The skill steps aside for small tasks and judgment-heavy debugging where delegation is not worthwhile.
- Stars
- ★ 1
- Forks
- ⑂ 0
- License
- MIT
Open prompt-injection-reviewGeneral skillSecurity and privacy
prompt-injection-review↗
@windchillscalanthes-ship-it
Traces prompt-injection paths from untrusted inputs to dangerous actions in LLM applications before release.
Project overviewprompt-injection-review maps trust boundaries in LLM applications that use RAG, external content, tools, agents, or model-rendered output, tracing untrusted sources through the model to outbound actions, execution, data access, or UI sinks. It recommends least privilege, human approval, egress controls, structured validation, and authorization enforced in code. It explicitly treats prompt injection as having no complete fix: the goal is to contain what a successful injection can reach, not to rely on filters for false certainty.
- Stars
- ★ 0
- Forks
- ⑂ 0
- License
- MIT
Open llm-cost-guardGeneral skillCode review and quality
llm-cost-guard↗
@windchillscalanthes-ship-it
Finds the dominant cost waste in LLM-calling code and proposes cheaper rewrites with explicit trade-offs.
Project overviewllm-cost-guard statically reviews LLM-calling code for repeated uncached prefixes, oversized model tiers, bloated context, unbounded output, unbudgeted agent loops, and repeated embeddings. It estimates the dominant tokens-times-price-times-volume term, then proposes prompt caching, model routing, context trimming, output caps, or batching while naming latency and quality trade-offs. It has no billing access, and projected savings still require current prices and evaluation results.
- Stars
- ★ 0
- Forks
- ⑂ 0
- License
- MIT
Open api-break-checkGeneral skillCode review and quality
api-break-check↗
@windchillscalanthes-ship-it
Finds backward-incompatible REST, GraphQL, and gRPC changes and proposes compatible alternatives.
Project overviewapi-break-check reviews API specifications, schemas, protobuf definitions, or implementation changes for removed response fields, newly required inputs, narrowed types, removed enum values, reused field numbers, and other breaking effects. It reasons about request-versus-response direction, explains which consumers fail, and proposes additive fields, deprecation windows, versioning, or reserved numbers. It is a pre-merge review rather than a substitute for contract-diff tools and direct consumer validation.
- Stars
- ★ 0
- Forks
- ⑂ 0
- License
- MIT
Open terraform-blast-radiusGeneral skillDevOps and deployment
terraform-blast-radius↗
@windchillscalanthes-ship-it
Catches destructive Terraform or OpenTofu changes before apply and proposes safer implementation paths.
Project overviewterraform-blast-radius reviews Terraform or OpenTofu plans and configuration diffs for forced replacements, resource renames, count-index shifts, dangerous state operations, and shared-resource cascades, distinguishing downtime from data-loss risk. It can propose prevent_destroy, create_before_destroy, moved blocks, for_each, or snapshot-based migrations without cloud access. Its output is a pre-apply review, not a guarantee, so the real plan, provider version, and tested backups still need checking.
- Stars
- ★ 0
- Forks
- ⑂ 1
- License
- MIT
Open n-plus-one-hunterGeneral skillCode review and quality
n-plus-one-hunter↗
@windchillscalanthes-ship-it
Finds ORM N+1 queries before merge and rewrites them with the appropriate loading strategy.
Project overviewn-plus-one-hunter statically reviews controllers, serializers, resolvers, and query code for classic or nested N+1s, per-row aggregates, cartesian eager-load blowups, and related patterns. It selects ORM-specific preload or join strategies for Rails, Django, Prisma, SQLAlchemy, and other listed stacks, explains the scaling cost, and reports a before-and-after query count. The proposed fix still needs verification against the actual ORM version, indexes, access pattern, and realistic data.
- Stars
- ★ 0
- Forks
- ⑂ 0
- License
- MIT
Open safe-migrationsGeneral skillDevOps and deployment
safe-migrations↗
@windchillscalanthes-ship-it
Review database migrations for locking and downtime risks
Project overviewsafe-migrations reviews PostgreSQL and MySQL schema changes for locks, table scans, and downtime risks before they merge. It explains the blocking mechanism, rewrites risky operations into copy-ready staged steps in SQL or framework idioms, and sequences expand-contract deployments when application changes must remain compatible.
- Stars
- ★ 0
- Forks
- ⑂ 0
- License
- MIT
Open ai-agent-skill-for-video-workflowCollection or directoryMedia, audio and video
ai-agent-skill-for-video-workflow↗
@dean9703111·Python
Convert audio into subtitles, refine them, and create social copy.
Project overviewThis Agent Skills collection covers a video subtitle workflow from audio transcription to publishing support. It converts common audio formats into SRT files, can refine wording and proper nouns against an `origin.md` transcript without changing timestamps, adds card annotations, and generates platform-specific summaries for Facebook, Threads, and YouTube. Each stage can be invoked separately or combined into the provided workflow. Users need to prepare the expected files in the same working directory and install the dependencies required by the individual skills.
- Stars
- ★ 195
- Forks
- ⑂ 46
- GitHub updated
- Jul 10, 2026
Open ARC-AGI-2Benchmark or evalResearch and knowledge
ARC-AGI-2↗
@arcprize
Train and evaluate systems on abstract grid-transformation reasoning tasks.
Project overviewARC-AGI-2 contains 1,000 public training tasks and 120 public evaluation tasks built from small colored-integer grids. Solvers infer a transformation from demonstration pairs and must produce every test output with the correct dimensions and cells. The evaluation set is intended for previously unseen testing, so repeatedly tuning against it would compromise the result.
- Stars
- ★ 723
- Forks
- ⑂ 106
- License
- Apache-2.0
Open Yagua-Image-ProcessorDesktop appImage generation and editingLinux
Yagua-Image-Processor↗
@GuilleBouix·Python
Batch-compress, convert, clean, and transform images on your desktop.
Project overviewYagua brings common image-processing jobs into a local desktop interface for designers, photographers, and web developers. It can batch-compress and convert files, resize or crop them, remove backgrounds, run OCR, vectorize images, and edit EXIF data without relying on a subscription web service. Windows and Linux are the main supported platforms, macOS remains experimental, and each module has its own practical batch limit.
- Stars
- ★ 165
- Forks
- ⑂ 20
- License
- MIT
Open academic-citation-audit-skillGeneral skillWriting and editing
academic-citation-audit-skill↗
@Chinelytra·Python
Audit manuscript references for authenticity, accuracy, consistency, and relevance.
Project overviewCheck an academic manuscript before submission for fabricated references, incorrect DOIs, bibliographic errors, missing text-to-list matches, and citations that do not support the associated claim. The skill accepts DOCX, TeX, and Bib inputs and uses CrossRef plus web search for cross-checking.
- Stars
- ★ 2
- Forks
- ⑂ 0
- GitHub updated
- Jul 10, 2026
Open report-writing-skillGeneral skillWriting and editing
report-writing-skill↗
@LostSunset·Python
Produce and validate Traditional Chinese technical reports for Taiwan delivery
Project overviewreport-writing is a reusable method library for formal DOCX reports, especially Traditional Chinese government and tender deliverables in Taiwan. It covers requirements, content strategy, layout, equations, clickable cross-references, page-numbered tables of contents, figures, and delivery checks, with linters for wording and English AI tells. The skill works with Claude Code and Codex and focuses on verified document output.
- Stars
- ★ 0
- Forks
- ⑂ 0
- License
- MIT