web-scraper-api
Production-grade web scraping with automatic anti-bot bypass, structured JSON parsing for 40+ targets, and geo-targeting. Use when the user needs to scrape web pages, extract product data, get search results, or collect structured data from supported e-commerce and search platfor
供給アセットの概要
リサーチとナレッジ作業
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
シナリオ
Document processing
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Agent 適合
Claude Code + Browser agents + CLI
Codex、Claude Code、Cursor、CLI、またはカスタム Agent に対応します。
インストール
準備完了
npx skills add oxylabs/agent-skills --skill web-scraper-api
メンテナンス
新しい
最終プッシュから 1 日
リスク
要レビュー
Permission surface may require sandboxing
GitHub 品質
566
74/100 品質 · 71/100 信頼
対象タグ
レビュー注記
Permission surface may require sandboxing · No explicit limitations or error handling guidance in SKILL.md (e.g., rate limits, timeouts, or failure scenarios).
Agent 導入スコアカード
信頼、監査、インストール準備状況を一目で確認
公開リポジトリのメタデータ、OpenAgentSkill のレビューシグナル、保守の鮮度、インストール準備状況を組み合わせたスコアです。候補選定の目安であり、人によるレビューの代替ではありません。
品質
強い本番ワークフローの候補に値する堅実な選択肢です。
信頼
サンドボックス限定信頼シグナルが不足または混在する有用な候補です。結果ループがタスク適合を示すまで、隔離されたワークスペースで使用してください。
監査
要レビューインストール準備、安全メタデータ、保守、採用リスクの機械可読なレビュー。
OpenAgentSkill Trust Score v5
インストール前に人のレビュー
実作業で使う前に、サンドボックスでのみ実行し、近い代替と比較してください。
スター
GitHub スター 566
リポジトリ活動
スター 566、フォーク 1
メンテナンス
最終プッシュから 1 日
ライセンス
MIT
インストール
npx skills add oxylabs/agent-skills --skill web-scraper-api
インストール安全性
標準パッケージまたはランタイムのインストールパス
権限範囲
shell or command execution, filesystem or document access
Agent の成果
Agent の成果データはまだありません
ドキュメント
README/SKILL.md の文脈が十分です
リスク概要
本番前にレビュー
- No explicit limitations or error handling guidance in SKILL.md (e.g., rate limits, timeouts, or failure scenarios).
- Quality score needs review
- Permission surface needs review: shell or command execution, filesystem or document access
- Stars/forks activity: 566 stars, 1 forks; issue activity unavailable in current metadata
インストール準備状況
インストールパスを利用可能
- インストールパスを利用できます
- リポジトリの根拠を利用できます
- ライセンスが明示されています
- Agent-Proven の成果エビデンスはまだありません
Agent 可読メタデータ
このスキルの機械可読な判断データ。
このブロックまたは埋め込み JSON を使い、Agent がこのスキルをインストールすべきか、代替を選ぶべきか、先に人のレビューを求めるべきかを判断できます。
適したタスク
- Web スクレイピング ワークフロー
- Claude Code チーム
- GitHub 採用シグナルを重視するチーム
- Crawl target URLs
適した Agent
インストール判断
- コマンド
- npx skills add oxylabs/agent-skills --skill web-scraper-api
- ポリシー
- ブロック
- 人によるレビュー
- はい
信頼とリスク
- 信頼
- 63/100
- 監査
- 79/100
- リスクレベル
- 要レビュー
成果ループ
- エンドポイント
- /api/agent/outcome
- イベント ID
- resolve
- 成果
- 5
インストールコマンド
npx skills add oxylabs/agent-skills --skill web-scraper-api使わない場合
- ベンダー提供の SLA が必要なチーム
- production agents without a repository review
- No explicit limitations or error handling guidance in SKILL.md (e.g., rate limits, timeouts, or failure scenarios).
- 高リスク権限のヒント: Shell or command execution, Secrets or environment access
- Permission surface may require sandboxing
Agent セーフティ v2
31/100 · 自動インストールを避ける
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
高
Shell またはコマンド実行
Skill メタデータに端末、CLI、Shell、サブプロセス、またはコマンド実行のワークフローが含まれます。
中
Browser automation
Skill may drive a browser or interact with web pages.
中
ネットワークアクセス
Skill はリモートページ、API、リポジトリ、外部サービスにアクセスする可能性があります。
中
ファイルシステムアクセス
Skill はプロジェクトファイル、ドキュメント、生成物、ローカルワークスペース状態を読み書きする可能性があります。
- 高リスク権限のヒント: Shell or command execution, Secrets or environment access
- Permission surface may require sandboxing
インストール先
Agent ワークフローにこのスキルをインストール
公開インストールエンドポイントからコマンド、安全チェックリスト、対象プロンプト、正規リンクを取得します。
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install oxylabs-web-scraper-apiAgent 解決プラン
インストール前に Agent に適合性を検証させます。
Resolve API は第一候補、代替、安全ポリシー、監査メモ、インストール先、Agent がそのまま使えるプロンプトを返します。
JSON を開く
/api/agent/resolve?task=Use%20web-scraper-api%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve テキスト
/api/agent/resolve?task=Use%20web-scraper-api%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
インストール引き継ぎ
/api/skills/oxylabs-web-scraper-api/install
Agent が確認すべきこと
- Resolve API でタスク適合と代替を確認。
- 監査・信頼スコアと安全ポリシーの警告を確認。
- Codex、Claude Code、Cursor、CLI のインストール先互換性を確認。
プロンプトをコピー
Task: Use web-scraper-api in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20web-scraper-api%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/oxylabs-web-scraper-api/install
Install command: npx skills add oxylabs/agent-skills --skill web-scraper-api
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent 引き継ぎ
別のディレクトリではなく、インストール経路を Agent に渡します。
公開インストールエンドポイントからコマンド、安全チェックリスト、対象プロンプト、正規リンクを取得します。
インストール引き継ぎ
/api/skills/oxylabs-web-scraper-api/install
LLM テキスト形式
/api/skills/oxylabs-web-scraper-api/install?format=text
代替を探す
/api/skills/search?q=web-scraper-api&limit=3
Agent プロンプト
Use web-scraper-api for this task. Review https://www.openagentskill.com/api/skills/oxylabs-web-scraper-api/install, then install with: npx skills add oxylabs/agent-skills --skill web-scraper-apiRegistry メタデータ
自動スキル選択用の Agent 可読プロファイル。
Registry API 経由で判断、信頼、監査、ユースケース、インストールのシグナルを提供し、UI をスクレイピングせずに Agent が順位付けできます。
Manifest
/api/registry/manifest/oxylabs-web-scraper-api
LLM テキスト
/api/registry/manifest/oxylabs-web-scraper-api?format=text
インストール別名
/api/registry/install/oxylabs-web-scraper-api
推奨
/api/registry/recommend?task=Use%20web-scraper-api%20in%20an%20agent%20workflow&limit=3
Agent 適合
Web スクレイピング
プラットフォーム
Claude Code, Browser agents
Agent 判断パネル
Web スクレイピング 向けの第一候補
有力候補として扱い、自分の Agent スタックで README とインストール経路を検証してください。
スタック内の役割
第一候補
主な適合
Web スクレイピング
信頼ラベル
本番対応
インストールパス
コマンド準備済み
使う場面
- Web スクレイピング ワークフロー
- Claude Code チーム
- GitHub 採用シグナルを重視するチーム
根拠
- GitHub スター 566
- 最近のリポジトリ活動
- インストールコマンドまたは GitHub リポジトリが利用可能
- 品質プロファイル 74/100
- OpenAgentSkill エンゲージメント 4 件
先にレビュー
- No explicit limitations or error handling guidance in SKILL.md (e.g., rate limits, timeouts, or failure scenarios).
実装パス
- 1サンドボックスの Agent にインストールし、Web スクレイピング タスクを一度最初から最後まで実行します。
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
信頼プロファイル
サンドボックス限定
信頼シグナルが不足または混在する有用な候補です。結果ループがタスク適合を示すまで、隔離されたワークスペースで使用してください。
GitHub 採用度
情報GitHub スター 566
スター/フォーク活動
確認スター 566、フォーク 1; 現在のメタデータでは Issue 活動を利用できません
最近のメンテナンス
合格最終プッシュから 1 日
ライセンスの明確さ
合格MIT
良いシグナル
- AI レビュー承認済み
- インストールパスを利用できます
- リポジトリの根拠を利用できます
- 最近保守されたリポジトリ
- 意味のある GitHub 採用シグナル
- インストールコマンドに明確な高リスクパターンはありません
- 成果ループは準備済みですが、最初の実行が必要です
インストール前にレビュー
- No explicit limitations or error handling guidance in SKILL.md (e.g., rate limits, timeouts, or failure scenarios).
- Quality score needs review
- Permission surface needs review: shell or command execution, filesystem or document access
- Stars/forks activity: 566 stars, 1 forks; issue activity unavailable in current metadata
- Permission surface: shell or command execution, filesystem or document access
- 実際の Agent 成果レポートはまだありません
- 無人インストールの前に人によるレビューが必要です
推奨アクション
実作業で使う前に、サンドボックスでのみ実行し、近い代替と比較してください。
品質プロファイル
強い Agent ワークフロー向けの候補
本番ワークフローの候補に値する堅実な選択肢です。
ワークフロー適合
このスキルを使うシナリオ
Collect structured data
Web scraping
I need my agent to scrape websites and extract structured data from pages.
Parse messy files
Document processing
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Operate local tools
Local desktop
I need my agent to operate local files and desktop apps in a repeatable workflow.
ワークフロー適合
完全なワークフローに追加
Scrape, clean, and reuse web data
Web data pipeline
A practical workflow for agents that crawl public pages, extract clean content, normalize data, and hand it to downstream research or RAG workflows.
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
代替候補
インストール前に比較
このタスクに適する可能性のある類似スキル。
Last30days Skill
Research the last 30 days across Reddit, X, YouTube, Hacker News, Polymarket, GitHub, and the web, then synthesize a grounded brief for an AI agent.
Academic Research Skills
Academic Research Skills for Claude Code: research → write → review → revise → finalize
GPT Researcher
Run autonomous deep research over web and local sources
DeepResearch
Tongyi Deep Research, the Leading Open-source Deep Research Agent
概要
--- name: web-scraper-api description: Production-grade web scraping with automatic anti-bot bypass, structured JSON parsing for 40+ targets, and geo-targeting. Use when the user needs to scrape web pages, extract product data, get search results, or collect structured data from supported e-commerce and search platforms without worrying about getting blocked and when geo targeting is required. ---
# Oxylabs Web Scraper API
## Authentication
Requires HTTP Basic Auth with credentials from environment variables:
```bash curl -u "$OXY_WSA_USERNAME:$OXY_WSA_PASSWORD" ... ```
## Endpoint
``` POST https://realtime.oxylabs.io/v1/queries # immediate response POST https://data.oxylabs.io/v1/queries # Push-Pull jobs, callbacks, storage Content-Type: application/json ```
## Core Parameters
| Parameter | Required | Description | |-----------|----------|-------------| | `source` | Yes | Target scraper (e.g., `universal`, `amazon_product`, `google_search`) | | `url` | Conditional | URL to scrape (for `universal` and `*_url` sources) | | `query` | Conditional | Search query or product ID (for `*_search` and `*_product` sources) | | `parse` | No | Enable structured data parsing (recommended for supported sources) | | `render` | No | JavaScript rendering: `html` or `png` | | `geo_location` | No | Geographic targeting: country/state/city, ZIP/postcode, coordinates, or Criteria ID where supported | | `session_id` | No | Reuse the same proxy IP across multiple jobs | | `content_encoding` | No | Set to `base64` when downloading image files via Realtime or Push-Pull | | `user_agent_type` | No | Device/browser preset, e.g., `desktop_chrome`, `mobile_ios`, `tablet_android` | | `locale` | No | Interface language / `Accept-Language`, e.g., `de-DE` | | `callback_url` | No | Push-Pull callback endpoint | | `storage_type`, `storage_url` | No | Push-Pull cloud upload target (`gcs`, `s3`, `tos`, `s3_compatible`) | | `markdown`, `xhr` | No | Enable markdown or captured XHR result types | | `browser_instructions` | No | Rendered browser actions; requires `render: "html"` | | `parsing_instructions`, `parser_preset` | No | Custom parser rules or saved preset; pair with `parse: true` | | `client_notes` | No | Client-side job tag saved with the job metadata | | `domain`, `subdomain`, `start_page`, `pages`, `limit`, `store_id`, `delivery_zip`, `fulfillment_type` | Source-specific | Marketplace/search/store localization and pagination fields |
`user_agent_type` values: `desktop`, `desktop_chrome`, `desktop_edge`, `desktop_firefox`, `desktop_opera`, `desktop_safari`, `mobile`, `mobile_android`, `mobile_ios`, `tablet`, `tablet_android`, `tablet_ios`.
## Context Parameters
Add these as `{ "key": "...", "value": ... }` objects in `context`:
| Key | Use | |-----|-----| | `force_headers`, `headers` | Merge custom headers with managed headers | | `force_cookies`, `cookies` | Merge custom cookies with managed cookies | | `http_method`, `content` | Use `post` with Base64-encoded body content | | `follow_redirects` | Follow 3xx redirect chains | | `successful_status_codes` | Treat specific non-standard HTTP codes as successful |
For multi-format output, enable types in the payload (`parse`, `markdown`, `xhr`, `render: "png"`) and request them with `?type=raw,parsed,png,markdown,xhr`.
For batch Push-Pull jobs, use `POST /v1/queries/batch` with arrays only for `query` or `url`; keep all other parameters singular. Maximum batch size is 5,000 values.
## Quick Start
**Scrape any URL:** ```bash curl -X POST 'https://realtime.oxylabs.io/v1/queries' \ -u "$OXY_WSA_USERNAME:$OXY_WSA_PASSWORD" \ -H 'Content-Type: application/json' \ -d '{"source": "universal", "url": "https://example.com"}' ```
**Google search with parsing:** ```bash curl -X POST 'https://realtime.oxylabs.io/v1/queries' \ -u "$OXY_WSA_USERNAME:$OXY_WSA_PASSWORD" \ -H 'Content-Type: application/json' \ -d '{"source": "google_search", "query": "best laptops", "parse": true}' ```
**Amazon product by ASIN:** ```bash curl -X POST 'https://realtime.oxylabs.io/v1/queries' \ -u "$OXY_WSA_USERNAME:$OXY_WSA_PASSWORD" \ -H 'Content-Type: application/json' \ -d '{"source": "amazon_product", "query": "B07FZ8S74R", "parse": true}' ```
## Choosing the Right Source
1. **Use specific sources when available** (`amazon_product`, `google_search`) - better parsing and reliability 2. **Use `universal` for unsupported sites** - works with any URL 3. **Enable `parse: true`** for structured JSON output on supported sources
## Response Structure
```json { "results": [{ "content": "...", "status_code": 200, "url": "https://..." }] } ```
With `parse: true`, `content` contains structured data (title, price, reviews, etc.) instead of raw HTML.
## Available Sources
For the complete list of 40+ supported sources organized by category, see [sources.md](sources.md).
## More Examples
For detailed request/response examples including geo-location, JavaScript rendering, and custom headers, see [examples.md](examples.md).
## Error Handling
| Code | Meaning | |------|---------| | 200 | Success | | 400 | Invalid parameters | | 401 | Authentication failed | | 403 | Access denied | | 429 | Rate limit exceeded |
## Key Guidelines
- Always set `parse: true` for supported sources to get structured data - Use ZIP codes for US e-commerce geo-location (e.g., `"90210"`) - Use country/state format for search engines (e.g., `"California,United States"`) - Add `render: "html"` for JavaScript-heavy pages - Use `render: ""` only to disable automatic forced rendering for force-rendered pages; set client timeouts near 180 seconds for rendered Realtime or Proxy Endpoint requests - Add `content_encoding: "base64"` when scraping image URLs, then decode `results[0].content` before saving the file
技術詳細
- バージョン
- 1.0.0
- ライセンス
- MIT
- 最終更新
- 2026年8月21日
- 公開日
- 2026年8月21日
判断の要約
第一候補
GitHub スター 566
Agent 実証エビデンス
Agent 実証エビデンス
Resolve、レビュー、インストール、限定実行後の成果レポート。
- 成功率
- —
- 直近の失敗
- —
- 成果
- 0
- 出力品質
- —
- 失敗
- 0
- 非該当
- 0
- インストール数
- 0
- リスクによりブロック
- 0
- 設定が必要
- 0
- 本番
- 0
Agent の実行結果はまだありません。最初の実行では /api/agent/outcome を通じて成功、設定要件、リスクによるブロック、失敗、非該当を報告できます。
成長ループ
共有キット
web-scraper-api 用のシナリオベース草案です。X へ手動投稿できます。
web-scraper-api: Production-grade web scraping with automatic anti-bot bypass, structured JSON parsing for 40+... 566 stars https://www.openagentskill.com/skills/oxylabs-web-scraper-api?ref=x
任意:インストールコマンド付きの返信
Listing + install path for web-scraper-api: https://www.openagentskill.com/skills/oxylabs-web-scraper-api?ref=x Install: npx skills add oxylabs/agent-skills --skill web-scraper-api
掲載元
Registry により登録
この掲載は公開ソースから登録されており、メンテナー申請が承認されるまで公式として表示されません。
- 作成者
- oxylabs
- インデックス作成者
- OpenAgentSkill コミュニティインデックス
帰属は公開リポジトリまたは作成者プロフィールにリンクされています。作成者は掲載を申請して所有権シグナルを更新できます。
このスキルを申請所有者の申請
このスキル掲載を申請
この Registry により登録 掲載は oxylabs に帰属していますが、まだ公式として表示されていません。申請すると、確認済み所有者シグナルが追加され、今後の公開、インストール、監査更新の信頼性が高まります。
クリエイター被リンクキット
README にエビデンスバッジを追加
開発者がリポジトリを評価する場所で、正規掲載、現在の信頼・監査シグナル、実際の Agent-Proven エビデンスを表示します。
[](https://www.openagentskill.com/skills/oxylabs-web-scraper-api)
[](https://www.openagentskill.com/skills/oxylabs-web-scraper-api)
[](https://www.openagentskill.com/skills/oxylabs-web-scraper-api/audit)
[](https://www.openagentskill.com/skills/oxylabs-web-scraper-api)作者
oxylabs
@oxylabs
プラットフォーム適合
健全性シグナル
- GitHub スター
- 566
- 品質スコア
- 42/100
- 最終 GitHub プッシュ
- 2026年8月21日
- フレームワークのヒント
- 不明
- OpenAgentSkill 閲覧数
- 4
- インストールコピー数
- 0
- 外部クリック
- 0
コミュニティシグナル
このスキルが Agent ワークフローに役立つかを共有してください。集約されたフィードバックがランキングを改善します。
信頼と安全性
サンドボックス限定
- GitHub 採用度GitHub スター 566情報
- スター/フォーク活動スター 566、フォーク 1; 現在のメタデータでは Issue 活動を利用できません確認
- 最近のメンテナンス最終プッシュから 1 日合格
- ライセンスの明確さMIT合格
- README/SKILL.md の完全性メタデータには十分な利用・ワークフロー文脈があります合格
- 依存関係/ランタイムのリスクcommand execution surface, network or browser surface情報
関連スキル
Last30days Skill
Research the last 30 days across Reddit, X, YouTube, Hacker News, Polymarket, GitHub, and the web, then synthesize a grounded brief for an AI agent.
53.5K スターAcademic Research Skills
Academic Research Skills for Claude Code: research → write → review → revise → finalize
38.4K スターGPT Researcher
Run autonomous deep research over web and local sources
28.0K スターDeepResearch
Tongyi Deep Research, the Leading Open-source Deep Research Agent
19.8K スター