Open SourceScraping17k
Crawlee
Production crawling library with queues, proxy rotation and sessions.
Crawlee is a web scraping and browser automation library from Apify for building reliable crawlers in Node.js/TypeScript and Python. It provides unified HTTP and headless-browser crawlers (Playwright, Puppeteer), automatic request queuing, proxy rotation, and anti-blocking features. It is notable for handling crawl-state management, scaling, and retries out of the box.
Repository
apify/crawleeLanguage
TypeScriptWhat you'd build with it
- Building large-scale crawlers with automatic queuing and retries
- Rotating proxies and managing sessions to avoid blocking
- Crawling JavaScript-heavy sites with managed headless browsers
Tags
crawlingscrapingproxy