web-scraper-api
Production-grade web scraping with automatic anti-bot bypass, structured JSON parsing for 40+ targets, and geo-targeting. Use when the user needs to scrape web pages, extract product data, get search results, or collect structured data from supported e-commerce and search platfor
供给资产档案
研究与知识工作
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
场景
Document processing
I need my agent to read PDFs, extract tables, and turn documents into structured data.
适配 Agent
Claude Code + Browser agents + CLI
适用于 Codex、Claude Code、Cursor、CLI 或自定义 Agent。
安装
就绪
npx skills add oxylabs/agent-skills --skill web-scraper-api
维护状态
新鲜
距上次推送 1 天
风险
需审查
Permission surface may require sandboxing
GitHub 质量
566
74/100 质量 · 71/100 信任
覆盖标签
审查说明
Permission surface may require sandboxing · No explicit limitations or error handling guidance in SKILL.md (e.g., rate limits, timeouts, or failure scenarios).
Agent 采用评分卡
一眼查看信任、审计与安装准备度
这些分数综合公开仓库元数据、OpenAgentSkill 审查信号、维护新鲜度与安装准备度。它用于候选筛选,不替代人工审查。
质量
强可靠的选择,值得加入生产工作流候选列表。
信任
仅限沙盒有用但信任信号不足或混杂的候选项。在结果闭环证明任务匹配前,请保持在隔离工作区内使用。
审计
需审查对安装准备度、安全元数据、维护情况与采用风险的机器可读审查。
OpenAgentSkill 信任评分 v5
安装前需人工审查
仅在沙盒中运行,并在用于真实工作前比较接近的替代方案。
Stars
566 个 GitHub Stars
仓库活跃度
566 个 Star,1 个 Fork
维护状态
距上次推送 1 天
许可证
MIT
安装
npx skills add oxylabs/agent-skills --skill web-scraper-api
安装安全性
标准软件包或运行时安装路径
权限范围
shell or command execution, filesystem or document access
Agent 结果
暂未有 Agent 结果数据
文档
README/SKILL.md 上下文充分
风险摘要
生产前审查
- No explicit limitations or error handling guidance in SKILL.md (e.g., rate limits, timeouts, or failure scenarios).
- Quality score needs review
- Permission surface needs review: shell or command execution, filesystem or document access
- Stars/forks activity: 566 stars, 1 forks; issue activity unavailable in current metadata
安装准备度
安装路径可用
- 安装路径可用
- 仓库证据可用
- 已声明许可证
- 暂无 Agent 验证结果证据
Agent 可读元数据
这个 Skill 的机器可读决策数据。
使用此区块或内嵌 JSON 判断 Agent 是否应安装该 Skill、选择替代方案,或先请求人工审查。
适用任务
- 网页抓取 工作流
- Claude Code 团队
- 重视 GitHub 采用信号的团队
- Crawl target URLs
适用 Agent
安装决策
- 命令
- npx skills add oxylabs/agent-skills --skill web-scraper-api
- 策略
- 阻止
- 人工审查
- 是
信任与风险
- 信任
- 63/100
- 审计
- 79/100
- 风险级别
- 需审查
结果闭环
- 端点
- /api/agent/outcome
- 事件 ID
- resolve
- 结果
- 5
不适用场景
- 需要厂商支持 SLA 的团队
- production agents without a repository review
- No explicit limitations or error handling guidance in SKILL.md (e.g., rate limits, timeouts, or failure scenarios).
- 高风险权限提示:Shell or command execution, Secrets or environment access
- Permission surface may require sandboxing
替代 Skill
Last30days Skill
53.5K Stars
npx skills add mvanhorn/last30days-skill -g
替代 Skill
Academic Research Skills
38.4K Stars
npx skills add Imbad0202/academic-research-skills
替代 Skill
GPT Researcher
28.0K Stars
npx skills add assafelovic/gpt-researcher
替代 Skill
DeepResearch
19.8K Stars
npx skills add Alibaba-NLP/DeepResearch
Agent 安全 v2
31/100 · 避免自动安装
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
高
Shell 或命令执行
Skill 元数据引用了终端、CLI、Shell、子进程或命令执行工作流。
中
Browser automation
Skill may drive a browser or interact with web pages.
中
网络访问
Skill 可能访问远程页面、API、仓库或外部服务。
中
文件系统访问
Skill 可能读取或写入项目文件、文档、生成产物或本地工作区状态。
- 高风险权限提示:Shell or command execution, Secrets or environment access
- Permission surface may require sandboxing
安装目标
在你的 Agent 工作流中安装此 Skill
通过公开安装端点获取命令、安全清单、目标提示词和该 Skill 的规范链接。
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install oxylabs-web-scraper-apiAgent 解析计划
让 Agent 在安装前验证匹配度。
Resolve API 返回首选 Skill、替代方案、安全策略、审计说明、安装目标和可直接执行的提示词,无需抓取此页面。
打开 JSON
/api/agent/resolve?task=Use%20web-scraper-api%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve 文本
/api/agent/resolve?task=Use%20web-scraper-api%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
安装交接
/api/skills/oxylabs-web-scraper-api/install
Agent 应检查
- 从 Resolve API 检查任务匹配与替代方案。
- 检查审计评分、信任评分和安全策略警告。
- 检查 Codex、Claude Code、Cursor 或 CLI 的安装目标兼容性。
复制提示词
Task: Use web-scraper-api in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20web-scraper-api%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/oxylabs-web-scraper-api/install
Install command: npx skills add oxylabs/agent-skills --skill web-scraper-api
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent 交接
把安装路径交给 Agent,而不是再给一个目录页。
通过公开安装端点获取命令、安全清单、目标提示词和该 Skill 的规范链接。
安装交接
/api/skills/oxylabs-web-scraper-api/install
LLM 文本格式
/api/skills/oxylabs-web-scraper-api/install?format=text
寻找替代方案
/api/skills/search?q=web-scraper-api&limit=3
Agent 提示词
Use web-scraper-api for this task. Review https://www.openagentskill.com/api/skills/oxylabs-web-scraper-api/install, then install with: npx skills add oxylabs/agent-skills --skill web-scraper-apiRegistry 元数据
用于自动选择 Skill 的 Agent 可读档案。
本页通过 Registry API 提供相同的决策、信任、审计、场景和安装信号,让 Agent 无需抓取界面即可排序。
Agent 决策面板
适合 网页抓取 的首选
将其作为优先候选,再在你的 Agent 环境中验证 README 与安装路径。
栈中角色
首选
主要匹配
网页抓取
信任标签
可用于生产
安装路径
命令已就绪
适用场景
- 网页抓取 工作流
- Claude Code 团队
- 重视 GitHub 采用信号的团队
证据
- 566 个 GitHub Stars
- 仓库近期活跃
- 已提供安装命令或 GitHub 仓库
- 74/100 质量档案
- 4 个 OpenAgentSkill 交互事件
先审查
- No explicit limitations or error handling guidance in SKILL.md (e.g., rate limits, timeouts, or failure scenarios).
实施路径
- 1在沙盒 Agent 中安装它,并端到端完成一次网页抓取任务。
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
信任档案
仅限沙盒
有用但信任信号不足或混杂的候选项。在结果闭环证明任务匹配前,请保持在隔离工作区内使用。
GitHub 采用度
信息566 个 GitHub Stars
Star/Fork 活跃度
检查566 个 Star,1 个 Fork; 当前元数据中没有议题活跃度信息
近期维护
通过距上次推送 1 天
许可证清晰度
通过MIT
积极信号
- AI 审查已通过
- 安装路径可用
- 仓库证据可用
- 近期维护的仓库
- 有意义的 GitHub 采用信号
- 安装命令未发现明显高风险模式
- 结果闭环已就绪,但需要首次真实 Agent 运行
安装前审查
- No explicit limitations or error handling guidance in SKILL.md (e.g., rate limits, timeouts, or failure scenarios).
- Quality score needs review
- Permission surface needs review: shell or command execution, filesystem or document access
- Stars/forks activity: 566 stars, 1 forks; issue activity unavailable in current metadata
- Permission surface: shell or command execution, filesystem or document access
- 暂未有真实 Agent 结果报告
- 无人值守安装前需要人工审查
建议操作
仅在沙盒中运行,并在用于真实工作前比较接近的替代方案。
质量档案
强 适用于 Agent 工作流的候选
可靠的选择,值得加入生产工作流候选列表。
工作流匹配
在这些场景使用此 Skill
Collect structured data
Web scraping
I need my agent to scrape websites and extract structured data from pages.
Parse messy files
Document processing
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Operate local tools
Local desktop
I need my agent to operate local files and desktop apps in a repeatable workflow.
工作流匹配
加入完整工作流
Scrape, clean, and reuse web data
Web data pipeline
A practical workflow for agents that crawl public pages, extract clean content, normalize data, and hand it to downstream research or RAG workflows.
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
替代方案短名单
安装前对比
可能适合该任务的相近 Skill。
Last30days Skill
Research the last 30 days across Reddit, X, YouTube, Hacker News, Polymarket, GitHub, and the web, then synthesize a grounded brief for an AI agent.
Academic Research Skills
Academic Research Skills for Claude Code: research → write → review → revise → finalize
GPT Researcher
Run autonomous deep research over web and local sources
DeepResearch
Tongyi Deep Research, the Leading Open-source Deep Research Agent
概览
--- name: web-scraper-api description: Production-grade web scraping with automatic anti-bot bypass, structured JSON parsing for 40+ targets, and geo-targeting. Use when the user needs to scrape web pages, extract product data, get search results, or collect structured data from supported e-commerce and search platforms without worrying about getting blocked and when geo targeting is required. ---
# Oxylabs Web Scraper API
## Authentication
Requires HTTP Basic Auth with credentials from environment variables:
```bash curl -u "$OXY_WSA_USERNAME:$OXY_WSA_PASSWORD" ... ```
## Endpoint
``` POST https://realtime.oxylabs.io/v1/queries # immediate response POST https://data.oxylabs.io/v1/queries # Push-Pull jobs, callbacks, storage Content-Type: application/json ```
## Core Parameters
| Parameter | Required | Description | |-----------|----------|-------------| | `source` | Yes | Target scraper (e.g., `universal`, `amazon_product`, `google_search`) | | `url` | Conditional | URL to scrape (for `universal` and `*_url` sources) | | `query` | Conditional | Search query or product ID (for `*_search` and `*_product` sources) | | `parse` | No | Enable structured data parsing (recommended for supported sources) | | `render` | No | JavaScript rendering: `html` or `png` | | `geo_location` | No | Geographic targeting: country/state/city, ZIP/postcode, coordinates, or Criteria ID where supported | | `session_id` | No | Reuse the same proxy IP across multiple jobs | | `content_encoding` | No | Set to `base64` when downloading image files via Realtime or Push-Pull | | `user_agent_type` | No | Device/browser preset, e.g., `desktop_chrome`, `mobile_ios`, `tablet_android` | | `locale` | No | Interface language / `Accept-Language`, e.g., `de-DE` | | `callback_url` | No | Push-Pull callback endpoint | | `storage_type`, `storage_url` | No | Push-Pull cloud upload target (`gcs`, `s3`, `tos`, `s3_compatible`) | | `markdown`, `xhr` | No | Enable markdown or captured XHR result types | | `browser_instructions` | No | Rendered browser actions; requires `render: "html"` | | `parsing_instructions`, `parser_preset` | No | Custom parser rules or saved preset; pair with `parse: true` | | `client_notes` | No | Client-side job tag saved with the job metadata | | `domain`, `subdomain`, `start_page`, `pages`, `limit`, `store_id`, `delivery_zip`, `fulfillment_type` | Source-specific | Marketplace/search/store localization and pagination fields |
`user_agent_type` values: `desktop`, `desktop_chrome`, `desktop_edge`, `desktop_firefox`, `desktop_opera`, `desktop_safari`, `mobile`, `mobile_android`, `mobile_ios`, `tablet`, `tablet_android`, `tablet_ios`.
## Context Parameters
Add these as `{ "key": "...", "value": ... }` objects in `context`:
| Key | Use | |-----|-----| | `force_headers`, `headers` | Merge custom headers with managed headers | | `force_cookies`, `cookies` | Merge custom cookies with managed cookies | | `http_method`, `content` | Use `post` with Base64-encoded body content | | `follow_redirects` | Follow 3xx redirect chains | | `successful_status_codes` | Treat specific non-standard HTTP codes as successful |
For multi-format output, enable types in the payload (`parse`, `markdown`, `xhr`, `render: "png"`) and request them with `?type=raw,parsed,png,markdown,xhr`.
For batch Push-Pull jobs, use `POST /v1/queries/batch` with arrays only for `query` or `url`; keep all other parameters singular. Maximum batch size is 5,000 values.
## Quick Start
**Scrape any URL:** ```bash curl -X POST 'https://realtime.oxylabs.io/v1/queries' \ -u "$OXY_WSA_USERNAME:$OXY_WSA_PASSWORD" \ -H 'Content-Type: application/json' \ -d '{"source": "universal", "url": "https://example.com"}' ```
**Google search with parsing:** ```bash curl -X POST 'https://realtime.oxylabs.io/v1/queries' \ -u "$OXY_WSA_USERNAME:$OXY_WSA_PASSWORD" \ -H 'Content-Type: application/json' \ -d '{"source": "google_search", "query": "best laptops", "parse": true}' ```
**Amazon product by ASIN:** ```bash curl -X POST 'https://realtime.oxylabs.io/v1/queries' \ -u "$OXY_WSA_USERNAME:$OXY_WSA_PASSWORD" \ -H 'Content-Type: application/json' \ -d '{"source": "amazon_product", "query": "B07FZ8S74R", "parse": true}' ```
## Choosing the Right Source
1. **Use specific sources when available** (`amazon_product`, `google_search`) - better parsing and reliability 2. **Use `universal` for unsupported sites** - works with any URL 3. **Enable `parse: true`** for structured JSON output on supported sources
## Response Structure
```json { "results": [{ "content": "...", "status_code": 200, "url": "https://..." }] } ```
With `parse: true`, `content` contains structured data (title, price, reviews, etc.) instead of raw HTML.
## Available Sources
For the complete list of 40+ supported sources organized by category, see [sources.md](sources.md).
## More Examples
For detailed request/response examples including geo-location, JavaScript rendering, and custom headers, see [examples.md](examples.md).
## Error Handling
| Code | Meaning | |------|---------| | 200 | Success | | 400 | Invalid parameters | | 401 | Authentication failed | | 403 | Access denied | | 429 | Rate limit exceeded |
## Key Guidelines
- Always set `parse: true` for supported sources to get structured data - Use ZIP codes for US e-commerce geo-location (e.g., `"90210"`) - Use country/state format for search engines (e.g., `"California,United States"`) - Add `render: "html"` for JavaScript-heavy pages - Use `render: ""` only to disable automatic forced rendering for force-rendered pages; set client timeouts near 180 seconds for rendered Realtime or Proxy Endpoint requests - Add `content_encoding: "base64"` when scraping image URLs, then decode `results[0].content` before saving the file
技术详情
- 版本
- 1.0.0
- 许可证
- MIT
- 最近更新
- 2026年8月21日
- 发布时间
- 2026年8月21日
决策摘要
首选
566 个 GitHub Stars
Agent 验证证据
Agent 验证证据
来自解析、审查、安装和一次小范围运行后的结果报告。
- 成功率
- —
- 近期失败
- —
- 结果
- 0
- 输出质量
- —
- 失败
- 0
- 不相关
- 0
- 安装次数
- 0
- 风险拦截
- 0
- 需要配置
- 0
- 生产环境
- 0
暂时没有 Agent 结果数据。首次 Agent 执行可以通过 /api/agent/outcome 报告成功、需要设置、风险拦截、失败或不相关。
增长闭环
分享工具包
为 web-scraper-api 准备的场景化草稿,可手动发布到 X。
web-scraper-api: Production-grade web scraping with automatic anti-bot bypass, structured JSON parsing for 40+... 566 stars https://www.openagentskill.com/skills/oxylabs-web-scraper-api?ref=x
可选:带安装命令的回复
Listing + install path for web-scraper-api: https://www.openagentskill.com/skills/oxylabs-web-scraper-api?ref=x Install: npx skills add oxylabs/agent-skills --skill web-scraper-api
收录来源
Registry 收录
此列表来自公开来源,维护者认领获批前不会标记为官方。
- 创作者
- oxylabs
- 收录方
- OpenAgentSkill 社区索引
归属链接指向公开仓库或创作者主页。创作者可认领列表以更新所有权信号。
认领此 Skill所有者认领
认领此 Skill 页面
这条 Registry 收录 列表归属于 oxylabs,但尚未标记为官方。认领后可增加已验证所有者信号,使后续发布、安装和审计更新更值得信赖。
创作者外链工具包
将证据徽章加入你的 README
在开发者评估仓库的位置展示规范页面、当前信任与审计信号,以及真实的 Agent 验证证据。
[](https://www.openagentskill.com/skills/oxylabs-web-scraper-api)
[](https://www.openagentskill.com/skills/oxylabs-web-scraper-api)
[](https://www.openagentskill.com/skills/oxylabs-web-scraper-api/audit)
[](https://www.openagentskill.com/skills/oxylabs-web-scraper-api)作者
oxylabs
@oxylabs
健康信号
- GitHub Stars
- 566
- 质量评分
- 42/100
- 最近 GitHub 推送
- 2026年8月21日
- 框架提示
- 未知
- OpenAgentSkill 浏览量
- 4
- 复制安装命令
- 0
- 跳转点击
- 0
社区信号
告诉我们这个 Skill 是否对你的 Agent 工作流有帮助。汇总反馈会持续改善排序。
信任与安全
仅限沙盒
- GitHub 采用度566 个 GitHub Stars信息
- Star/Fork 活跃度566 个 Star,1 个 Fork; 当前元数据中没有议题活跃度信息检查
- 近期维护距上次推送 1 天通过
- 许可证清晰度MIT通过
- README/SKILL.md 完整度元数据包含足够的用法与工作流上下文通过
- 依赖与运行时风险command execution surface, network or browser surface信息
相关 Skill
Last30days Skill
Research the last 30 days across Reddit, X, YouTube, Hacker News, Polymarket, GitHub, and the web, then synthesize a grounded brief for an AI agent.
53.5K StarsAcademic Research Skills
Academic Research Skills for Claude Code: research → write → review → revise → finalize
38.4K StarsGPT Researcher
Run autonomous deep research over web and local sources
28.0K StarsDeepResearch
Tongyi Deep Research, the Leading Open-source Deep Research Agent
19.8K Stars