MinerU
opendatalab
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 311
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 311
opendatalab
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
ossappscollective
Document scanning app
dataelement
BISHENG is an open LLM devops platform for next generation Enterprise AI applications. Powerful and comprehensive features include: GenAI workflow, RAG, Agent, Unified m…
datalab-to
OCR model that handles complex tables, forms, handwriting with full layout.
Unstructured-IO
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for languag…
JaidedAI
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
forthespada
🔥🔥超过1000本的计算机经典书籍、个人笔记资料以及本人在各平台发表文章中所涉及的资源等。书籍资源包括C/C++、Java、Python、Go语言、数据结构与算法、操作系统、后端架构、计算机系统知识、数据库、计算机网络、设计模式、前端、汇编以及校招社招各种面经~
pymupdf
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
papermark
Papermark is the open-source DocSend alternative with built-in analytics and custom domains.
axa-group
Transforms PDF, Documents and Images into Enriched Structured Data
microsoft
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
PaddlePaddle
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ lan…
windingwind
Translate PDF, EPub, webpage, metadata, annotations, notes to the target language. Support 20+ translate services.