Project overview
Indic Voice Pipeline combines yt-dlp with fine-tuned Whisper models and Qwen3-ASR to process audio or video in 12 Indian languages and export text, subtitles, and structured JSON. Optional pyannote diarization adds speaker labels and requires a Hugging Face token. Transcription runs locally, though model downloads and hardware demands can be substantial.
Repository facts
- Primary language
- Python
- License
- MIT
- Repository updated
- Feb 15, 2026
- Default branch
- main
Resource types
General skill
Use cases
Media, audio and videoContent creation
Platforms
Claude Code, Codex, and more
Runtime
Local
Capabilities
Audio processingTranscriptionTranslationVideo processing
Audience
Content creators