The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns
Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
RAG and knowledge
I need my agent to build a RAG workflow over documents and retrieve reliable context.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add ArchiveTeam/grab-site
Maintenance
stale
1y since push
Risk
Needs review
License is unclear
GitHub quality
1.6K
67/100 quality · 78/100 trust
Coverage tags
Review notes
License is unclear · Repository appears stale
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
PromisingUseful candidate, but compare it with alternatives before adopting.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewInstall readiness, security metadata, maintenance, and adoption risk.
Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
1.6K GitHub stars
Repo activity
1.6K stars, 156 forks
Maintenance
1y since push
License
Unknown
Install
npx skills add ArchiveTeam/grab-site
Install safety
standard package or runtime install path
Permission surface
filesystem or document access, network or browser access
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add ArchiveTeam/grab-siteDo not use when
Agent safety v2
Sparse or mixed signals. Useful for discovery, but not for autonomous installation.
Test manually in an isolated workspace and compare against safer alternatives.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
Install targets
Copy the registry command or an agent-specific install prompt for Codex, Claude Code, and Cursor.
Use the registry command when your workflow supports the OpenAgentSkill installer.
$ npx skills add ArchiveTeam/grab-siteAgent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Resolve JSON
/api/agent/resolve?task=Use%20Grab%20Site%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20Grab%20Site%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/archiveteam-grab-site/install
Agent should check
Copy prompt
Task: Use Grab Site in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20Grab%20Site%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/archiveteam-grab-site/install
Install command: npx skills add ArchiveTeam/grab-site
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/archiveteam-grab-site/install
LLM text format
/api/skills/archiveteam-grab-site/install?format=text
Find alternatives
/api/skills/search?q=Grab%20Site&limit=3
Agent prompt
Use Grab Site for this task. Review https://www.openagentskill.com/api/skills/archiveteam-grab-site/install, then install with: npx skills add ArchiveTeam/grab-siteRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/archiveteam-grab-site
LLM text
/api/registry/manifest/archiveteam-grab-site?format=text
Install alias
/api/registry/install/archiveteam-grab-site
Recommend
/api/registry/recommend?task=Use%20Grab%20Site%20in%20an%20agent%20workflow&limit=3
Agent fit
Web scraping
Use-case tags
Platforms
Python, Crawler, Claude Code
Audit report
Review install readiness, maintenance, trust, quality, and metadata warnings before adding this skill to an agent workflow.
Agent decision cockpit
Prototype with this skill first; keep a fallback candidate ready.
Role in stack
Fallback candidate
Primary fit
Web scraping
Trust label
Prototype first
Install path
Command ready
Use when
Evidence
Review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
PASS1.6K GitHub stars
Stars/forks activity
INFO1.6K stars, 156 forks; issue activity unavailable in current metadata
Recent maintenance
FIX1y since push
License clarity
CHECKUnknown
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Useful candidate, but compare it with alternatives before adopting.
Workflow fit
Collect structured data
I need my agent to scrape websites and extract structured data from pages.
Analyze matches
I need my agent to analyze football matches, World Cup data, xG, players, teams, and predictions.
Build and ship code
I need a coding agent that can understand a repository, edit code, and review pull requests.
Stack fit
Scrape, clean, and reuse web data
A practical stack for agents that crawl public pages, extract clean content, normalize data, and hand it to downstream research or RAG workflows.
Turn skills into distribution
A stack for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Inspect, patch, and verify code
A stack for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Alternative shortlist
Similar skills in this category, ranked with the same readiness and quality signals.
Open-source LLM-friendly web crawler and scraper
Adaptive web scraping for agent data collection
Scrapy, a fast high-level web crawling & scraping framework for Python.
A visual no-code/code-free web crawler/spider易采集:一个可视化浏览器自动化测试/数据采集/网页爬虫软件,可以无代码图形化的设计和执行爬虫任务。别名:ServiceWrapper面向Web应用的智能化服务封装系统。
The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns
Imported by the skill-only GitHub discovery pipeline because it matches agent skill, automation, domain workflow, RAG, document-processing, data, finance, security, or developer-tool signals. Protocol-server projects are excluded from automated imports.
Frameworks & Tools
Decision snapshot
1,584 GitHub stars
Audit snapshot
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the audit before production use.
Growth loop
Scenario-led draft for Grab Site, ready for a manual X post.
Most web agents fail in the boring part: messy pages, missing context, repeatable extraction. Grab Site gives agents a cleaner path to browse, extract, and monitor web pages. 1.6K stars https://www.openagentskill.com/skills/archiveteam-grab-site?ref=x #AIAgents
Listing + install path for Grab Site: https://www.openagentskill.com/skills/archiveteam-grab-site?ref=x Install: npx skills add ArchiveTeam/grab-site
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This community indexed listing is attributed to ArchiveTeam but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/archiveteam-grab-site)
[](https://www.openagentskill.com/skills/archiveteam-grab-site)
[](https://www.openagentskill.com/skills/archiveteam-grab-site/audit)
[](https://www.openagentskill.com/skills/archiveteam-grab-site)ArchiveTeam✓
@archiveteam
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
Crawl4AI
Open-source LLM-friendly web crawler and scraper
73.1K stars · 31.0K installsScrapling
Adaptive web scraping for agent data collection
70.0K stars · 0 installsScrapy
Scrapy, a fast high-level web crawling & scraping framework for Python.
62.5K stars · 0 installsEasySpider
A visual no-code/code-free web crawler/spider易采集:一个可视化浏览器自动化测试/数据采集/网页爬虫软件,可以无代码图形化的设计和执行爬虫任务。别名:ServiceWrapper面向Web应用的智能化服务封装系统。
44.1K stars · 0 installs