工作流方案

网页数据管道 工作流

为「网页数据管道」设计的可审查 Agent 工作流。

为「网页数据管道」设计的可审查 Agent 工作流。

预期结果

  • 明确「网页数据管道」的目标
  • 选择匹配的能力
  • 在进入下一步前核验结果

工作流地图

按此顺序执行

01

界定任务

将目标转成清晰、可审查的 Agent 任务。

02

选择能力

为当前阶段选择最匹配的能力,而不是浏览泛化目录。

03

执行并检查

在来源、权限和范围已核对的前提下完成最小有效步骤。

04

核验结论

在进入下一阶段前审查证据和结果。

推荐能力

为每个步骤选择技能

按当前工作流相关性、质量、GitHub 采用度与维护活跃度排序。这是决策指南,不是单一安装命令。

#1Crawlee质量 · 100

源仓库说明(原文)

Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.

24K StarsApache-2.0browser-automation
对比
$ npx skills add apify/crawlee
#2Firecrawl质量 · 100

源仓库说明(原文)

The API to search, scrape, and interact with the web at scale. 🔥

139K StarsAGPL-3.0agent-frameworks
对比
$ npx skills add firecrawl/firecrawl
#3Crawlee Python质量 · 100

源仓库说明(原文)

Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.

9.2K StarsApache-2.0browser-automation
对比
$ npx skills add apify/crawlee-python
#4Scrapegraph AI质量 · 100

源仓库说明(原文)

Python scraper based on AI

27K StarsMITweb-automation
对比
$ npx skills add ScrapeGraphAI/Scrapegraph-ai
#5Maxun质量 · 100

源仓库说明(原文)

🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥

16K StarsAGPL-3.0web-automation
对比
$ npx skills add getmaxun/maxun
#6Crawl4AI质量 · 100

源仓库说明(原文)

Open-source LLM-friendly web crawler and scraper

73K StarsApache-2.0web-automation
对比
$ npx skills add unclecode/crawl4ai
#7Lux质量 · 100

源仓库说明(原文)

👾 Fast and simple video download library and CLI tool written in Go

31K StarsMITweb-automation
对比
$ npx skills add iawia002/lux
#8Colly质量 · 100

源仓库说明(原文)

Elegant Scraper and Crawler Framework for Golang

25K StarsApache-2.0web-automation
对比
$ npx skills add gocolly/colly

适用场景

  • - 网页数据管道 工作流
  • - 需要留存审计依据的 Agent 决策
  • - 可控工作区内的试运行

以下情况不适合使用

  • - 任务需要无人监督地访问敏感系统
  • - 无法审查来源材料或结果质量

需要可直接执行的技能包?

技能包包含安装顺序、审计链接和机器可读的 Agent 计划。

浏览技能包