General skillMedia, audio and video

scenelens

Extract scene-aware frames, OCR text, and chunked transcripts from videos.

Stars
3
Forks
0
License
MIT
Updated
Updated May 8, 2026

Checking live repository facts…

Project overview

SceneLens prepares public video URLs or local files for Claude by extracting frames at scene changes, running OCR on each frame, and attaching a timestamped transcript. It first tries native captions, then falls back to Groq or OpenAI Whisper and automatically splits audio that exceeds the API size limit before merging timestamps. Accuracy is best for videos under ten minutes, output is capped at 100 frames, and authenticated private platforms are not supported.

Repository facts

Primary language
Python
License
MIT
Repository updated
May 8, 2026
Default branch
main

Resource types

General skill

Use cases

Media, audio and video

Platforms

Claude Code, Codex, and more

Runtime

Local

Capabilities

Transcription

Related projects

Browse more similar projects