MinerU
opendatalab
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 64
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 64
opendatalab
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
pymupdf
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
datalab-to
OCR model that handles complex tables, forms, handwriting with full layout.
PaddlePaddle
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ lan…
Python-Markdown
A Python implementation of John Gruber’s Markdown with Extension support.
JaidedAI
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
Textualize
Rich is a Python library for rich text and beautiful formatting in the terminal.
ocrmypdf
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
hiroi-sora
OCR software, free and offline. 开源、免费的离线OCR软件。支持截屏/批量导入图片,PDF文档识别,排除水印/页眉页脚,扫描/生成二维码。内置多国语言库。
py-pdf
A pure-python PDF library capable of splitting, merging, cropping, and transforming the pages of PDF files
Textualize
Rich-cli is a command line toolbox for fancy output in the terminal
mwouts
Jupyter Notebooks as Markdown Documents, Julia, Python or R scripts