Annuaire de skills

Découvrez des skills réutilisables pour les AI agents.

Recherchez de vrais skills GitHub par tâche et vérifiez Stars, confiance, audit, catégorie et chemin d’installation avant de les utiliser.

Chaque recommandation reste clairement reliée à son dépôt, son audit et son chemin d’installation.

Résultats de recherche: crawlers

Annuaire en anglais

Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.

24K
Stars
85/100
Confiance
Catégorie: browser-automationAudit

Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.

9.2K
Stars
83/100
Confiance
Catégorie: browser-automationAudit

一些非常有趣的python爬虫例子,对新手比较友好,主要爬取淘宝、天猫、微信、微信读书、豆瓣、QQ等网站。(Some interesting examples of python crawlers that are friendly to beginners. )

15K
Stars
76/100
Confiance
Catégorie: web-automationAudit

Norconex Crawlers (or spiders) are flexible web and filesystem crawlers for collecting, parsing, and manipulating data from the web or filesystem to various data repositories such as search engines.

202
Stars
68/100
Confiance
Catégorie: rag-knowledgeAudit

A multi-thread crawler framework with many builtin image crawlers provided.

930
Stars
66/100
Confiance
Catégorie: web-automationAudit

Detailed web scraping tutorials for dummies with financial data crawlers on Reddit WallStreetBets, CME (both options and futures), US Treasury, CFTC, LME, MacroTrends, SHFE and alternative data crawlers on Tomtom, BBC, Wall Street Journal, Al Jazeera, Reuters, Financial Times, Bloomberg, CNN, Fortune, The Economist

884
Stars
65/100
Confiance
Catégorie: web-automationAudit

Serve clean Markdown from your Next.js site to AI agents, crawlers, and LLMs. Humans get HTML, agents get clean Markdown of the same pages. Two-file install.

58
Stars
67/100
Confiance
Catégorie: agent-frameworksAudit

Whole-repo audits in eight modes. Codebase — merged structural + correctness audit, should this exist AND does it do what it promises. Triggers "nuclear review", "code judo", "whole codebase review", "should this exist", "adversarial audit", "fable audit", "correctness audit", "expectation gaps". Docs/Process — doc drift, walkable journeys. Triggers "audit the docs", "doc drift", "process audit", "walk the journeys". Performance — measured-only perf audit; no finding without a number. Triggers "perf audit", "performance audit", "why is it slow", "bundle audit", "build is slow". Threat-model — abuse paths. Triggers "threat model", "STRIDE", "attack surface". Motion — animation audit. Triggers "motion audit", "audit the animations". SEO — discoverability + AEO. Triggers "seo audit", "aeo", "answer engine", "llms.txt", "rank better". Debt — `SHORTCUT:` ledger. Triggers "debt ledger", "shortcut ledger". Owns bare "audit the codebase"; single-page CWV fix loops go to /lighthouse.

42
Stars
60/100
Confiance
Catégorie: securityAudit