Anthropic PDF
Read, extract, OCR and assemble PDFs
Setup: The source's Python / CLI dependencies; OCR needs an OCR engine
Scans require OCR. Table structure and reading order must be checked against page images.
Read pinned SKILL.md ↗Practical task guide
Choose a PDF skill for Claude Code by task: digital text, tables, scanned-page OCR or markdown conversion. Check dependencies and reconcile the extracted output.
Decision prompt
Help me read a digital report. Ask for missing inputs, choose by output, explain dependencies, and deliver a checked file without inventing results.
Published by OpenAgentSkill. This workflow follows pinned upstream instructions. Inspect the source, prerequisites and deliverable checks before installing. Editorial update: .
Task fit, source instructions, setup and limitations.
Read, extract, OCR and assemble PDFs
Recommended shortlist
Read, extract, OCR and assemble PDFs
Setup: The source's Python / CLI dependencies; OCR needs an OCR engine
Scans require OCR. Table structure and reading order must be checked against page images.
Read pinned SKILL.md ↗How to use this guide
Inspect a representative PDF page and determine whether text is selectable.
Use the source's digital extraction tools, or its OCR path for scanned images.
Keep page numbers and source references with the extracted content; do not silently replace missing text.
Reconcile table rows and numeric totals and check reading order before writing markdown or importing into RAG.
Evaluation notes
First check whether the PDF contains selectable text. Digital extraction, scanned-page OCR and table reconstruction need different operations and checks. Visual PDF creation is outside this parsing shortlist.
Source review, an editorial example and a successful runtime test answer different questions. Record which one you actually have.
FAQ
The page may be an image scan rather than digital text. Check the page visually and use OCR when needed. Extraction can also fail on unusual encodings or protected files.
Start with extraction that fits the input, verify its reading order and tables, then convert the result. The PDF workflow is the primary source here; a converter library is a dependency, not automatically an installable agent skill.
No. It describes the pinned source workflow and the checks required. This release does not include a comparative OCR runtime test.
Next guides
Practical task guide
Compare frontend design skills by task: landing pages, product dashboards, Figma implementation, accessibility review and React performance. Explore editable examples.
Practical task guide
Understand official Remotion agent skills, their setup and limitations. Download an original React composition and MP4 example, and follow the render verification steps.
Practical task guide
Use an XLSX skill for CSV cleanup, typed cells, formulas and editable Excel delivery. Download a messy CSV and the checked workbook example with missing values preserved.