web-scraper-api
Production-grade web scraping with automatic anti-bot bypass, structured JSON parsing for 40+ targets, and geo-targeting. Use when the user needs to scrape web pages, extract product data, get search results, or collect structured data from supported e-commerce and search platfor
Supply asset profile
Research and knowledge work
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Document processing
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Agent fit
Claude Code + Browser agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add oxylabs/agent-skills --skill web-scraper-api
Maintenance
fresh
1d since push
Risk
Needs review
Permission surface may require sandboxing
GitHub quality
566
74/100 Quality · 71/100 Trust
Coverage tags
Review notes
Permission surface may require sandboxing · No explicit limitations or error handling guidance in SKILL.md (e.g., rate limits, timeouts, or failure scenarios).
Agent adoption scorecard
Trust, audit, and install readiness at a glance
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Human review before install
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
566 GitHub stars
Repo activity
566 stars, 1 forks
Maintenance
1d since push
License
MIT
Install
npx skills add oxylabs/agent-skills --skill web-scraper-api
Install safety
standard package or runtime install path
Permission surface
shell or command execution, filesystem or document access
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Review before production
- No explicit limitations or error handling guidance in SKILL.md (e.g., rate limits, timeouts, or failure scenarios).
- Quality score needs review
- Permission surface needs review: shell or command execution, filesystem or document access
- Stars/forks activity: 566 stars, 1 forks; issue activity unavailable in current metadata
Install readiness
Install path available
- Install path is available
- Repository evidence is available
- License is declared
- No Agent Proven outcome evidence yet
Agent-readable metadata
Machine-readable decision data for this skill.
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
- Web scraping workflows
- Claude Code teams
- teams that value GitHub adoption signals
- Crawl target URLs
Suited agents
Install decision
- Command
- npx skills add oxylabs/agent-skills --skill web-scraper-api
- Policy
- block
- Human review
- yes
Trust and risk
- Trust
- 63/100
- Audit
- 79/100
- Risk level
- Needs review
Outcome loop
- Endpoint
- /api/agent/outcome
- Event ID
- resolve
- Outcomes
- 5
Install command
npx skills add oxylabs/agent-skills --skill web-scraper-apiDo not use when
- teams that need a vendor-supported SLA
- production agents without a repository review
- No explicit limitations or error handling guidance in SKILL.md (e.g., rate limits, timeouts, or failure scenarios).
- High-risk permission hints: Shell or command execution, Secrets or environment access
- Permission surface may require sandboxing
Alternative
Last30days Skill
53.5K Stars
npx skills add mvanhorn/last30days-skill -g
Alternative
Academic Research Skills
38.4K Stars
npx skills add Imbad0202/academic-research-skills
Alternative
GPT Researcher
28.0K Stars
npx skills add assafelovic/gpt-researcher
Alternative
DeepResearch
19.8K Stars
npx skills add Alibaba-NLP/DeepResearch
Agent safety v2
31/100 · Avoid automatic install
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
high
Shell or command execution
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Browser automation
Skill may drive a browser or interact with web pages.
medium
Network access
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Filesystem access
Skill may read or write project files, documents, generated artifacts, or local workspace state.
- High-risk permission hints: Shell or command execution, Secrets or environment access
- Permission surface may require sandboxing
Install targets
Install this skill in your agent workflow
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install oxylabs-web-scraper-apiAgent resolve plan
Let an agent verify fit before installing.
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20web-scraper-api%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20web-scraper-api%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/oxylabs-web-scraper-api/install
Agent should check
- Task fit and alternatives from Resolve API.
- Audit score, trust score, and safety policy warnings.
- Install target compatibility for Codex, Claude Code, Cursor, or CLI.
Copy prompt
Task: Use web-scraper-api in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20web-scraper-api%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/oxylabs-web-scraper-api/install
Install command: npx skills add oxylabs/agent-skills --skill web-scraper-api
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Give an agent the install path, not another directory page.
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/oxylabs-web-scraper-api/install
LLM text format
/api/skills/oxylabs-web-scraper-api/install?format=text
Find alternatives
/api/skills/search?q=web-scraper-api&limit=3
Agent prompt
Use web-scraper-api for this task. Review https://www.openagentskill.com/api/skills/oxylabs-web-scraper-api/install, then install with: npx skills add oxylabs/agent-skills --skill web-scraper-apiRegistry metadata
Agent-readable profile for automatic skill selection.
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/oxylabs-web-scraper-api
LLM text
/api/registry/manifest/oxylabs-web-scraper-api?format=text
Install alias
/api/registry/install/oxylabs-web-scraper-api
Recommend
/api/registry/recommend?task=Use%20web-scraper-api%20in%20an%20agent%20workflow&limit=3
Agent fit
Web scraping
Use-case tags
Platforms
Claude Code, Browser agents
Audit report
Needs review · 79/100
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Primary pick for Web scraping
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
Web scraping
Trust label
Production-ready
Install path
Command ready
Use when
- Web scraping workflows
- Claude Code teams
- teams that value GitHub adoption signals
Evidence
- 566 GitHub stars
- recent repository activity
- install command or GitHub repo available
- 74/100 quality profile
- 4 OpenAgentSkill engagement events
review first
- No explicit limitations or error handling guidance in SKILL.md (e.g., rate limits, timeouts, or failure scenarios).
Implementation path
- 1Install it in a sandbox agent and run one Web scraping task end to end.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
Trust profile
Sandbox only
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
INFO566 GitHub stars
Stars/forks activity
CHECK566 stars, 1 forks; issue activity unavailable in current metadata
Recent maintenance
PASS1d since push
License clarity
PASSMIT
Good signals
- AI review approved
- Install path is available
- Repository evidence is available
- Recently maintained repository
- Meaningful GitHub adoption signal
- Install command has no obvious high-risk pattern
- Outcome loop is ready but needs first real agent run
Review before install
- No explicit limitations or error handling guidance in SKILL.md (e.g., rate limits, timeouts, or failure scenarios).
- Quality score needs review
- Permission surface needs review: shell or command execution, filesystem or document access
- Stars/forks activity: 566 stars, 1 forks; issue activity unavailable in current metadata
- Permission surface: shell or command execution, filesystem or document access
- No real agent outcome reports yet
- Human review required before unattended installation
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Strong candidate for agent workflows
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Use this skill in these scenarios
Collect structured data
Web scraping
I need my agent to scrape websites and extract structured data from pages.
Parse messy files
Document processing
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Operate local tools
Local desktop
I need my agent to operate local files and desktop apps in a repeatable workflow.
Workflow fit
Add it to a complete workflow
Scrape, clean, and reuse web data
Web data pipeline
A practical workflow for agents that crawl public pages, extract clean content, normalize data, and hand it to downstream research or RAG workflows.
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Alternative shortlist
Compare before you install
Similar skills that may fit this task.
Last30days Skill
Research the last 30 days across Reddit, X, YouTube, Hacker News, Polymarket, GitHub, and the web, then synthesize a grounded brief for an AI agent.
Academic Research Skills
Academic Research Skills for Claude Code: research → write → review → revise → finalize
GPT Researcher
Run autonomous deep research over web and local sources
DeepResearch
Tongyi Deep Research, the Leading Open-source Deep Research Agent
Overview
--- name: web-scraper-api description: Production-grade web scraping with automatic anti-bot bypass, structured JSON parsing for 40+ targets, and geo-targeting. Use when the user needs to scrape web pages, extract product data, get search results, or collect structured data from supported e-commerce and search platforms without worrying about getting blocked and when geo targeting is required. ---
# Oxylabs Web Scraper API
## Authentication
Requires HTTP Basic Auth with credentials from environment variables:
```bash curl -u "$OXY_WSA_USERNAME:$OXY_WSA_PASSWORD" ... ```
## Endpoint
``` POST https://realtime.oxylabs.io/v1/queries # immediate response POST https://data.oxylabs.io/v1/queries # Push-Pull jobs, callbacks, storage Content-Type: application/json ```
## Core Parameters
| Parameter | Required | Description | |-----------|----------|-------------| | `source` | Yes | Target scraper (e.g., `universal`, `amazon_product`, `google_search`) | | `url` | Conditional | URL to scrape (for `universal` and `*_url` sources) | | `query` | Conditional | Search query or product ID (for `*_search` and `*_product` sources) | | `parse` | No | Enable structured data parsing (recommended for supported sources) | | `render` | No | JavaScript rendering: `html` or `png` | | `geo_location` | No | Geographic targeting: country/state/city, ZIP/postcode, coordinates, or Criteria ID where supported | | `session_id` | No | Reuse the same proxy IP across multiple jobs | | `content_encoding` | No | Set to `base64` when downloading image files via Realtime or Push-Pull | | `user_agent_type` | No | Device/browser preset, e.g., `desktop_chrome`, `mobile_ios`, `tablet_android` | | `locale` | No | Interface language / `Accept-Language`, e.g., `de-DE` | | `callback_url` | No | Push-Pull callback endpoint | | `storage_type`, `storage_url` | No | Push-Pull cloud upload target (`gcs`, `s3`, `tos`, `s3_compatible`) | | `markdown`, `xhr` | No | Enable markdown or captured XHR result types | | `browser_instructions` | No | Rendered browser actions; requires `render: "html"` | | `parsing_instructions`, `parser_preset` | No | Custom parser rules or saved preset; pair with `parse: true` | | `client_notes` | No | Client-side job tag saved with the job metadata | | `domain`, `subdomain`, `start_page`, `pages`, `limit`, `store_id`, `delivery_zip`, `fulfillment_type` | Source-specific | Marketplace/search/store localization and pagination fields |
`user_agent_type` values: `desktop`, `desktop_chrome`, `desktop_edge`, `desktop_firefox`, `desktop_opera`, `desktop_safari`, `mobile`, `mobile_android`, `mobile_ios`, `tablet`, `tablet_android`, `tablet_ios`.
## Context Parameters
Add these as `{ "key": "...", "value": ... }` objects in `context`:
| Key | Use | |-----|-----| | `force_headers`, `headers` | Merge custom headers with managed headers | | `force_cookies`, `cookies` | Merge custom cookies with managed cookies | | `http_method`, `content` | Use `post` with Base64-encoded body content | | `follow_redirects` | Follow 3xx redirect chains | | `successful_status_codes` | Treat specific non-standard HTTP codes as successful |
For multi-format output, enable types in the payload (`parse`, `markdown`, `xhr`, `render: "png"`) and request them with `?type=raw,parsed,png,markdown,xhr`.
For batch Push-Pull jobs, use `POST /v1/queries/batch` with arrays only for `query` or `url`; keep all other parameters singular. Maximum batch size is 5,000 values.
## Quick Start
**Scrape any URL:** ```bash curl -X POST 'https://realtime.oxylabs.io/v1/queries' \ -u "$OXY_WSA_USERNAME:$OXY_WSA_PASSWORD" \ -H 'Content-Type: application/json' \ -d '{"source": "universal", "url": "https://example.com"}' ```
**Google search with parsing:** ```bash curl -X POST 'https://realtime.oxylabs.io/v1/queries' \ -u "$OXY_WSA_USERNAME:$OXY_WSA_PASSWORD" \ -H 'Content-Type: application/json' \ -d '{"source": "google_search", "query": "best laptops", "parse": true}' ```
**Amazon product by ASIN:** ```bash curl -X POST 'https://realtime.oxylabs.io/v1/queries' \ -u "$OXY_WSA_USERNAME:$OXY_WSA_PASSWORD" \ -H 'Content-Type: application/json' \ -d '{"source": "amazon_product", "query": "B07FZ8S74R", "parse": true}' ```
## Choosing the Right Source
1. **Use specific sources when available** (`amazon_product`, `google_search`) - better parsing and reliability 2. **Use `universal` for unsupported sites** - works with any URL 3. **Enable `parse: true`** for structured JSON output on supported sources
## Response Structure
```json { "results": [{ "content": "...", "status_code": 200, "url": "https://..." }] } ```
With `parse: true`, `content` contains structured data (title, price, reviews, etc.) instead of raw HTML.
## Available Sources
For the complete list of 40+ supported sources organized by category, see [sources.md](sources.md).
## More Examples
For detailed request/response examples including geo-location, JavaScript rendering, and custom headers, see [examples.md](examples.md).
## Error Handling
| Code | Meaning | |------|---------| | 200 | Success | | 400 | Invalid parameters | | 401 | Authentication failed | | 403 | Access denied | | 429 | Rate limit exceeded |
## Key Guidelines
- Always set `parse: true` for supported sources to get structured data - Use ZIP codes for US e-commerce geo-location (e.g., `"90210"`) - Use country/state format for search engines (e.g., `"California,United States"`) - Add `render: "html"` for JavaScript-heavy pages - Use `render: ""` only to disable automatic forced rendering for force-rendered pages; set client timeouts near 180 seconds for rendered Realtime or Proxy Endpoint requests - Add `content_encoding: "base64"` when scraping image URLs, then decode `results[0].content` before saving the file
Technical details
- Version
- 1.0.0
- License
- MIT
- Last updated
- Aug 21, 2026
- Published
- Aug 21, 2026
Decision snapshot
Primary pick
566 GitHub stars
Audit
Install review
Install and adoption review
- Security
- 76/100
- Maintenance
- 100/100
- Install
- 92/100
Agent-proven evidence
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
- Success rate
- —
- Recent failure
- —
- Outcomes
- 0
- Output quality
- —
- Failed
- 0
- Not relevant
- 0
- Installs
- 0
- Risk blocked
- 0
- Setup needed
- 0
- Production
- 0
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Add to agent workflow
Free and open source. Review the report before installing into production agents.
Growth loop
Share kit
Scenario-led draft for web-scraper-api, ready for a manual X post.
web-scraper-api: Production-grade web scraping with automatic anti-bot bypass, structured JSON parsing for 40+... 566 stars https://www.openagentskill.com/skills/oxylabs-web-scraper-api?ref=x
Optional reply with install command
Listing + install path for web-scraper-api: https://www.openagentskill.com/skills/oxylabs-web-scraper-api?ref=x Install: npx skills add oxylabs/agent-skills --skill web-scraper-api
Listing source
Registry indexed
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
- Creator
- oxylabs
- Source
- oxylabs/agent-skills
- Indexed by
- OpenAgentSkill community index
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
Claim this skill listing
This Registry indexed listing is attributed to oxylabs but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Add the evidence badges to your README
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/oxylabs-web-scraper-api)
[](https://www.openagentskill.com/skills/oxylabs-web-scraper-api)
[](https://www.openagentskill.com/skills/oxylabs-web-scraper-api/audit)
[](https://www.openagentskill.com/skills/oxylabs-web-scraper-api)Author
oxylabs
@oxylabs
Tags
Platform fit
Health signals
- GitHub stars
- 566
- Quality score
- 42/100
- Last GitHub push
- Aug 21, 2026
- Framework hints
- Unknown
- OpenAgentSkill views
- 4
- Install copies
- 0
- Outbound clicks
- 0
Community signal
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Trust & safety
Sandbox only
- GitHub adoption566 GitHub starsINFO
- Stars/forks activity566 stars, 1 forks; issue activity unavailable in current metadataCHECK
- Recent maintenance1d since pushPASS
- License clarityMITPASS
- README/SKILL.md completenessMetadata includes enough usage and workflow contextPASS
- Dependency/runtime riskcommand execution surface, network or browser surfaceINFO
Related skills
Last30days Skill
Research the last 30 days across Reddit, X, YouTube, Hacker News, Polymarket, GitHub, and the web, then synthesize a grounded brief for an AI agent.
53.5K StarsAcademic Research Skills
Academic Research Skills for Claude Code: research → write → review → revise → finalize
38.4K StarsGPT Researcher
Run autonomous deep research over web and local sources
28.0K StarsDeepResearch
Tongyi Deep Research, the Leading Open-source Deep Research Agent
19.8K Stars