scenelens
Extract scene-aware frames, OCR text, and chunked transcripts from videos.
Extract scene-aware frames, OCR text, and chunked transcripts from videos.
SceneLens prepares public video URLs or local files for Claude by extracting frames at scene changes, running OCR on each frame, and attaching a timestamped transcript. It first tries native captions, then falls back to Groq or OpenAI Whisper and automatically splits audio that exceeds the API size limit before merging timestamps. Accuracy is best for videos under ten minutes, output is capped at 100 frames, and authenticated private platforms are not supported.
Resource types
Use cases
Platforms
Runtime
Turn character references into prompts for game gameplay videos.
Download audio from one Bilibili video or an uploader's full catalog.
Capabilities
Public GitHub facts last synced Jul 10, 2026.
Animate routes, locations, and regions into editorial map videos
Create finished AI videos across multiple models from one Claude Code workflow.