MinerU
opendatalab
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 36
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 36
opendatalab
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Textualize
Rich is a Python library for rich text and beautiful formatting in the terminal.
Unstructured-IO
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for languag…
QuestPDF
QuestPDF is a modern library for PDF document generation. Its fluent C# API lets you design complex layouts with clean, readable code. Create documents using a flexible,…
bytedance
The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
embedpdf
A PDF viewer that seamlessly integrates with any JavaScript project
unidoc
Golang PDF library for creating and processing PDF files (pure go)
commonmark
Java library for parsing and rendering CommonMark (Markdown)
UglyToad
Read and extract text and other content from PDFs in C# (port of PDFBox)
jingsongliujing
基于PaddleOCR重构,并且脱离PaddlePaddle深度学习训练框架的轻量级OCR,推理速度超快 —— A lightweight OCR system based on PaddleOCR, decoupled from the PaddlePaddle deep learning training framework, wi…
rust-skia
Rust Bindings for the Skia Graphics Library
kotaro-kinoshita
YomiTokuはAIを活用した日本語文書解析エンジンを提供するPythonパッケージです。 Yomitoku is an AI-powered document image analysis package designed specifically for the Japanese language.
Jaspersoft
JasperReports® - Free Java Reporting Library