General skillTesting and debugging

gpu-perf-playbook

Diagnose GPU bottlenecks in PyTorch training and LLM inference workloads.

Stars
1
Forks
0
License
MIT
Updated
Updated Apr 11, 2026

Checking live repository facts…

Project overview

Use the playbook directly or install its Claude skill to investigate slow training, low utilization, data-loading stalls, out-of-memory failures, weak multi-GPU scaling, and LLM serving latency or throughput problems. Its workflow starts with measurement, classifies the limiting resource, applies the smallest relevant change, and verifies the result. Separate guides cover training, inference, distributed systems, kernels, profilers, and recurring performance anti-patterns in code review.

Repository facts

Primary language
Shell
License
MIT
Repository updated
Apr 11, 2026
Default branch
main

Resource types

General skill

Use cases

Testing and debugging

Platforms

Claude Code, Codex, and more

Capabilities

Verification and evals

Related projects

Browse more similar projects