PyMuPDF
pymupdf
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Search retrieves candidates across the registry and ranks a bounded shortlist by task fit. This count is matching candidates, not the registry total. No suitable match? Try a specific tool or task.
1β16 / 32
Results: 32
pymupdf
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
firecrawl
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
adbar
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
landing-ai
This tool has been deprecated. Use Agentic Document Extraction instead.
landing-ai
Python library for Agentic Document Extraction (ADE).
codelucas
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
getmaxun
π₯ The open-source no-code platform for web scraping, crawling, search and AI data extraction β’ Turn websites into structured APIs in minutes π₯
Achno
A tool to convert a Wallpaper's color scheme / palette, OCR with VLM's Traditional & Hybrid, Image Compression ,color palette extraction, image upsacling with Adversariaβ¦
NanoNets
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
unjs
π PDF extraction and rendering across all JavaScript runtimes
soxoj
βοΈ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
AndyTheFactory
π° Newspaper4k a fork of the beloved Newspaper3k. Extraction of articles, titles, and metadata from news websites.
openai
Automate a real browser from the terminal for navigation, form filling, snapshots, screenshots, extraction, and UI-flow debugging.
nexu-io
A two-spread digital e-guide preview β page 1 is a cover (display title, author, "What's inside" stats, table of contents teaser); page 2 is a spread (lesson body with pβ¦
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.