Project overview
Convert manuals, datasheets, programming guides, and other structured PDFs with a digital-first extraction path and optional OCR for scanned pages. The skill reconstructs headings, recovers tables and preformatted listings, exports figures, and supports page ranges or batch folders. It writes Markdown beside the source by default, requires uv and Python, and can use pdftotext as an optional geometry source for better layout recovery.
Repository facts
- Primary language
- Python
- License
- MIT
- Repository updated
- Apr 13, 2026
- Default branch
- main
Resource types
General skill
Use cases
Document productivityResearch and knowledge
Platforms
Claude Code, Codex, and more
Runtime
Command line
Audience
Developers