Project overview
scraper offers an article mode for clean text and metadata and a soup mode for CSS-selected fields. It respects robots.txt, handles retry and exponential backoff, writes JSONL incrementally, resumes successful URLs after interruption, and can seed work from a sitemap. The tool identifies HTML, PDF, and JSON responses and can optionally use ScraperAPI for proxying or JavaScript rendering. It is intended for auditable collection rather than unbounded crawling.
Repository facts
- Primary language
- HTML
- License
- MIT
- Repository updated
- Jun 21, 2026
- Default branch
- main
Resource types
General skill
Use cases
Research and knowledge
Platforms
Claude Code, Codex, and more