Trafilatura
adbar
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
Page 1 · 16 shown · 71 public entries
Results: 71
adbar
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
papermark
Papermark is the open-source DocSend alternative with built-in analytics and custom domains.
4paradigm
OpenMLDB is an open-source machine learning database that provides a feature platform computing consistent features for training and inference.
forthespada
🔥🔥超过1000本的计算机经典书籍、个人笔记资料以及本人在各平台发表文章中所涉及的资源等。书籍资源包括C/C++、Java、Python、Go语言、数据结构与算法、操作系统、后端架构、计算机系统知识、数据库、计算机网络、设计模式、前端、汇编以及校招社招各种面经~
extract internal monitoring data from application logs for collection in a timeseries database
JosephLai241
Universal Reddit Scraper - A comprehensive Reddit scraping/archival command-line tool.
alvarobartt
Financial Data Extraction from Investing.com with Python
diskoverdata
Diskover Community Edition - Open source file indexer, file search engine and data management and analytics powered by Elasticsearch
501351981
支持word(.docx)、excel(.xlsx,.xls)、pdf、pptx等各类型office文件预览的vue组件集合,提供一站式office文件预览方案,支持vue2和3,也支持React等非Vue框架。Web-based pdf, excel, word, pptx preview library
opsdisk
pagodo (Passive Google Dork) - Automate Google Hacking Database scraping and searching
guyueyingmu
AV 电影管理系统, avmoo , javbus , javlibrary 爬虫,线上 AV 影片图书馆,AV 磁力链接数据库,Japanese Adult Video Library,Adult Video Magnet Links - Japanese Adult Video Database
tabulapdf
Tabula is a tool for liberating data tables trapped inside PDF files
NanoNets
Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extract…
sec-edgar
Download all companies periodic reports, filings and forms from EDGAR database.
SpiderClub
:zap: A distributed crawler for weibo, building with celery and requests.
SkywalkerDarren
ChatWeb can crawl web pages, read PDF, DOCX, TXT, and extract the main content, then answer your questions based on the content, or summarize the key points.