Project overview
SceneLens prepares public video URLs or local files for Claude by extracting frames at scene changes, running OCR on each frame, and attaching a timestamped transcript. It first tries native captions, then falls back to Groq or OpenAI Whisper and automatically splits audio that exceeds the API size limit before merging timestamps. Accuracy is best for videos under ten minutes, output is capped at 100 frames, and authenticated private platforms are not supported.
Repository facts
- Primary language
- Python
- License
- MIT
- Repository updated
- May 8, 2026
- Default branch
- main
Resource types
General skill
Use cases
Media, audio and video
Platforms
Claude Code, Codex, and more
Runtime
Local
Capabilities
Transcription