Newspaper
codelucas
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–16 / 29
Results: 29
codelucas
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
bda-research
Web Crawler/Spider for NodeJS + server-side jQuery ;-)
hiddendevj
Collection of China illegal cases about web crawler 本项目用来整理所有中国大陆爬虫开发者涉诉与违规相关的新闻、资料与法律法规。致力于帮助在中国大陆工作的爬虫行业从业者了解我国相关法律,避免触碰数据合规红线。
BruceDone
A collection of awesome web crawler,spider in different languages
apache
A scalable, mature and versatile web crawler based on Apache Storm
fredwu
A high performance web crawler / scraper in Elixir.
xuxueli
A lightweight web crawler framework.(Java爬虫框架)
hect0x7
Python API for JMComic | 提供Python API访问禁漫天堂,同时支持网页端和移动端 | 禁漫天堂GitHub Actions下载器🚀
dataabc
新浪微博爬虫,用python爬取新浪微博数据,并下载微博图片和微博视频
JayBizzle
🕷 CrawlerDetect is a PHP class for detecting bots/crawlers/spiders via the user agent
dadoonet
Elasticsearch File System Crawler (FS Crawler)
webrecorder
Run a high-fidelity browser-based web archiving crawler in a single Docker container
dixudx
Easily download all the photos/videos from tumblr blogs. 下载指定的 Tumblr 博客中的图片,视频
palewire
An open-source archive that gathers, saves, shares and analyzes news homepages
oxylabs
Crawl a website starting from a URL, find relevant pages, and extract data – all guided by your natural language prompt.
code4craft
A scalable web crawler framework for Java.