Document skills

PDF extraction skills for AI agents.

Compare skills for PDF parsing, OCR, table extraction, markdown conversion, document metadata, and agent-ready file processing.

Built for users searching for AI agent skills that can parse PDFs, extract tables, and convert documents into usable context.

Matched

5

Stars

385K

Input

PDF

Output

Markdown

Agent jobs

Start from a real workflow, not a keyword.

These pages are built for high-intent search and for agents that need a structured shortlist with install commands, trust signals, audit links, and real outcome evidence before installing third-party code.

01

Convert PDFs into clean markdown for agents

02

Extract tables and metadata from reports

03

Prepare legal, finance, and research documents for review

04

Use OCR fallback when scanned pages need text extraction

Ranked shortlist

High-signal skills to inspect first.

Open best list
80K stars

Convert PDFs, Office documents, and web files into clean markdown for agents.

100

Quality

79

Trust

—

Proven

Document ProcessingJun 1, 2026 pushMITNeeds first agent run
$ npx skills add microsoft/markitdown
163K stars

Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.

100

Quality

97

Trust

77

Proven

design-creativeJul 17, 2026 pushSource terms (see LICENSE.txt)Promising agent evidence
$ npx skills add anthropics/skills --skill canvas-design
34K stars

Turn websites into clean markdown or structured data for retrieval and agents.

100

Quality

80

Trust

—

Proven

Web ScrapingJun 1, 2026 pushAGPL-3.0Needs first agent run
$ npx skills add mendableai/firecrawl
66K stars

Open-source LLM-friendly web crawler and scraper for agent workflows.

100

Quality

79

Trust

39

Proven

Web ScrapingJun 1, 2026 pushApache-2.0Early agent signal
$ npx skills add unclecode/crawl4ai
42K stars

Data framework for building RAG and knowledge workflows around agent tasks.

100

Quality

79

Trust

—

Proven

RAGJun 1, 2026 pushMITNeeds first agent run
$ npx skills add run-llama/llama_index

Evaluation

How to choose the right skill.

Handles layout, headings, and tables without destroying context

Reports extraction limits and OCR uncertainty

Supports batch or repeatable processing

Documents privacy and local processing assumptions

Questions

Which PDF skill should I choose first?

Choose a skill that supports your document type, preserves tables or headings, and makes extraction failures visible instead of silently guessing.

Can these skills handle scanned PDFs?

Some can, but scanned PDFs usually need OCR and human review for high-stakes documents.