技能目录

为 AI Agent 发现可复用技能。

按任务搜索真实的 GitHub 技能,并在使用前查看 Stars、信任、审计、分类和安装路径。

每个推荐都保留与其仓库、审计和安装路径的明确关联。

搜索结果: scraping

英文目录

Adaptive web scraping for agent data collection

70K
Stars
75/100
信任
分类: web-automation审计

Scrapy, a fast high-level web crawling & scraping framework for Python.

63K
Stars
82/100
信任
分类: web-automation审计

Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.

24K
Stars
85/100
信任
分类: browser-automation审计

🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥

16K
Stars
85/100
信任
分类: web-automation审计

The headless browser for AI agents and web scraping

16K
Stars
80/100
信任
分类: web-automation审计

SeleniumBase is a framework for UI Testing, Web Scraping, and Stealth. Passes every bot-detection test with CDP Mode, and extends Playwright.

13K
Stars
85/100
信任
分类: testing-qa审计

Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML

6.4K
Stars
85/100
信任
分类: web-automation审计

Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.

9.2K
Stars
83/100
信任
分类: browser-automation审计

An AI-powered research assistant that performs iterative, deep research on any topic by combining search engines, web scraping, and large language models. The goal of this repo is to provide the simplest implementation of a deep research agent - e.g. an agent that can refine its research direction overtime and deep dive into a topic.

19K
Stars
85/100
信任
分类: research审计

List of libraries, tools and APIs for web scraping and data processing.

7.9K
Stars
76/100
信任
分类: web-automation审计
Rod81

A Chrome DevTools Protocol driver for web automation and scraping.

7.0K
Stars
81/100
信任
分类: web-automation审计

Stealth headless browser for AI agents — bypass Cloudflare, bot detection, and anti-scraping. Drop-in Puppeteer/Playwright replacement.

6.9K
Stars
84/100
信任
分类: agent-frameworks审计