Project overview
This skill transcribes spoken video with Whisper word timestamps, groups words into 2–5-word chunks around punctuation and pauses, and burns styled captions with ffmpeg. You can choose fonts, colors, brand plates, aspect-ratio presets, languages, and a glossary; it requires a Whisper package and ffmpeg.
Repository facts
- Primary language
- Not detected
- License
- MIT
- Repository updated
- May 19, 2026
- Default branch
- main
Resource types
General skill
Use cases
Video creation and editing
Platforms
Claude Code, Codex, and more
Capabilities
TranscriptionVideo processing