OpenAgentSkill Task Task: Crawl a documentation site Intent: Turn documentation pages into clean markdown or records that an agent can search and reuse. Agent prompt: Find the best skill for crawling a documentation website and converting pages into clean markdown with useful metadata. Success criteria: - Preserves source URLs - Produces clean markdown - Can limit crawl scope Do not use when: - Docs block crawling - The content is private without authorization - You need pixel-perfect browser state Resolve API: https://www.openagentskill.com/api/agent/resolve?task=Find%20the%20best%20skill%20for%20crawling%20a%20documentation%20website%20and%20converting%20pages%20into%20clean%20markdown%20with%20useful%20metadata.&agent=codex&max_risk=medium Ranked skills: 1. AnyCrawl (any4ai-anycrawl) Match score: 744 AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured SERP results from Google/Bing/Baidu/etc. Native multi-threading for bulk processing. Trust: 86/100 Production candidate Audit: 91/100 Safe to try Install: Detail: https://www.openagentskill.com/skills/any4ai-anycrawl Install API: https://www.openagentskill.com/api/skills/any4ai-anycrawl/install --- 2. Trafilatura (adbar-trafilatura) Match score: 715 Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML Trust: 86/100 Production candidate Audit: 91/100 Safe to try Install: Detail: https://www.openagentskill.com/skills/adbar-trafilatura Install API: https://www.openagentskill.com/api/skills/adbar-trafilatura/install --- 3. Firecrawl (firecrawl) Match score: 690 Turn any website into LLM-ready markdown or structured data Trust: 87/100 Production candidate Audit: 90/100 Safe to try Install: Detail: https://www.openagentskill.com/skills/firecrawl Install API: https://www.openagentskill.com/api/skills/firecrawl/install --- 4. Crawlee (apify-crawlee) Match score: 689 Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation. Trust: 89/100 Production candidate Audit: 93/100 Safe to try Install: Detail: https://www.openagentskill.com/skills/apify-crawlee Install API: https://www.openagentskill.com/api/skills/apify-crawlee/install --- 5. Crawlee Python (apify-crawlee-python) Match score: 683 Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation. Trust: 88/100 Production candidate Audit: 93/100 Safe to try Install: Detail: https://www.openagentskill.com/skills/apify-crawlee-python Install API: https://www.openagentskill.com/api/skills/apify-crawlee-python/install