
crawlee-python
apify
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
AI 简介
Crawlee-Python 是一个面向 Python 开发者的网页爬虫与浏览器自动化库,用于构建高可靠性、抗反爬的网络数据采集系统。它支持 HTTP 请求、Playwright 无头/有头浏览器、Parsel 和 BeautifulSoup 解析器,并内置代理轮换、请求调度、数据持久化及自动重试机制,可高效提取 HTML、PDF、图片等各类网页资源,适配 AI 数据准备(如 RAG、LLM 训练)和企业级数据采集场景。
Python
Apache License 2.09.3k
Stars
769
Forks
46
Watchers
74
Issues
Star 增长
今日0
近 7 天0
近 30 天+23
综合评分66.96
默认分支master