Docext
NanoNets
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 48
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 48
NanoNets
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
fhamborg
news-please - an integrated web crawler and information extractor for news that just works
landing-ai
This tool has been deprecated. Use Agentic Document Extraction instead.
landing-ai
Python library for Agentic Document Extraction (ADE).
jsvine
Plumb a PDF for detailed information about each char, rectangle, line, et cetera — and easily extract text and tables.
dynobo
OCR powered screen-capture tool to capture information instead of images
Autumn-27
ScopeSentry-Cyberspace mapping, subdomain enumeration, port scanning, sensitive information discovery, vulnerability scanning, distributed nodes
codelucas
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
getmaxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
pymupdf
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
getomni-ai
OCR & Document Extraction using vision models
megadose
holehe allows you to check if the mail is used on different sites like twitter, instagram and will retrieve information on sites with the forgotten password function.
Achno
A tool to convert a Wallpaper's color scheme / palette, OCR with VLM's Traditional & Hybrid, Image Compression ,color palette extraction, image upsacling with Adversaria…
firecrawl
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.