Docext
NanoNets
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1β16 / 160
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 160
NanoNets
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
JaidedAI
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
landing-ai
This tool has been deprecated. Use Agentic Document Extraction instead.
landing-ai
Python library for Agentic Document Extraction (ADE).
PaddlePaddle
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ lanβ¦
jsvine
Plumb a PDF for detailed information about each char, rectangle, line, et cetera β and easily extract text and tables.
iamgio
πͺ Markdown with superpowers: from ideas to papers, presentations, websites, books, and knowledge bases.
Stirling-Tools
#1 PDF Application on GitHub that lets you edit PDFs on any device anywhere
tesseract-ocr
Tesseract Open Source OCR Engine (main repository)
opendatalab
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
docling-project
Get your documents ready for gen AI
paperless-ngx
A community-supported supercharged document management system: scan, index and archive all your documents
ShareX
ShareX is a free and open-source application that enables users to capture or record any area of their screen with a single keystroke. It also supports uploading images,β¦
ocrmypdf
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
dynobo
OCR powered screen-capture tool to capture information instead of images