Scrapy, a fast high-level web crawling & scraping framework for Python.
Web scraping - 精选合集
Web scraping
The web is the world's largest dataset. These open source projects help developers crawl, extract, and run scrapers at scale.
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, an…
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Wo…
Elegant Scraper and Crawler Framework for Golang
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Get web data for AI agents and LLMs
A next-generation crawling and spidering framework.
The context API to search, scrape, and interact with the web at scale. 🔥
🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
Python scraper based on AI
Playwright is a framework for Web Testing and Automation. It allows testing Chromium, Firefox and WebKit with a single API.
JavaScript API for Chrome and Firefox
A browser automation framework and ecosystem.
📊 APIs for web automation, testing, and bypassing bot-detection.
Lightpanda: the headless browser designed for AI and automation
Deploy headless browsers in Docker. Run on our cloud or bring your own. Free for non-commercial uses.
The fast, flexible, and elegant library for parsing and manipulating HTML and XML.
jsoup: the Java HTML parser, built for HTML editing, cleaning, scraping, and XSS safety.
Nokogiri (鋸) makes it easy and painless to work with XML and HTML from Ruby.
A little like that j-thing, only in Go.