PaddleOCR
PaddlePaddle
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ lan…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
Page 1 · 16 shown · 17 public entries
Results: 17
PaddlePaddle
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ lan…
opendataloader-project
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
dataelement
BISHENG is an open LLM devops platform for next generation Enterprise AI applications. Powerful and comprehensive features include: GenAI workflow, RAG, Agent, Unified m…
Sumanth077
A curated collection of practical AI projects implementing OCR systems, RAG, AI agents, and other AI use cases.
opendataloader-project
Procedure for extracting structured data from PDFs with opendataloader-pdf (ODL) correctly: read the installed tool's own help to discover its current options, build the…
yfedoseev
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders,…
yobix-ai
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
anbeime
软件开发工程师与数据科学家在构建RAG系统时,当需处理PDF/Word/Excel等多格式复杂文档,用此技能可自动触发三级解析降级,一键输出高置信度结构化Markdown与元数据,免去繁琐清洗,直接夯实企业知识库数据底座!
bzsanti
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
NoEdgeAI
A python wrapper for the Doc2X API and comes with native texts processing (to improve PDF recall in RAG). | Doc2X API的python封装,同时附带本地的文本处理(提升PDF在RAG中的召回率)。
Extends pageIndex into an AI document workspace with multi-format parsing, OCR, visual TOC, custom models, citations, and agentic QA.
AKSarav
PDFStract - Extract, Chunking and Embedding Layer in Your RAG Pipeline - Available as CLI - WEBUI - API
mrmps
Browser based tool to convert PDFs to Markdown
KylinMountain
Convert files into markdown to help RAG or LLM understand, based on markitdown and MinerU, which could provide high quality pdf parser.
nico-martin
A Webapp that uses Retrieval Augmented Generation (RAG) and Large Language Models to interact with a PDF directly in the browser.
mehmetba
mehmetba/pdf-analyze-streamlit is a high-star GitHub project relevant to AI agent workflows.