An iOS OCR Server Using Apple’s Vision Framework
$ npx skills add riddleling/iOS-OCR-ServerScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
AI Agent Skill Repository
Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.
Live registry search
Exact name and slug matches are checked against the live registry before ranked alternatives.
Decision filters
Showing 1-16 of 120 ranked candidates matching "macos-vision-ocr"
Best blend of relevance, quality, freshness, and verified outcomes
An iOS OCR Server Using Apple’s Vision Framework
$ npx skills add riddleling/iOS-OCR-ServerScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
Rust library and CLI tool for OCR (extracting text from images)
$ npx skills add robertknight/ocrsScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
让纯文本模型更好地做视觉任务的DeepSeek Harness插件:带意图的图片问答、长截图 OCR、UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, g…
$ npx skills add Anionex/dsh-vision-toolkitScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
Claude Code + CLI · 4 targets
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ lan…
$ npx skills add PaddlePaddle/PaddleOCRScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
OCR software, free and offline. 开源、免费的离线OCR软件。支持截屏/批量导入图片,PDF文档识别,排除水印/页眉页脚,扫描/生成二维码。内置多国语言库。
$ npx skills add hiroi-sora/Umi-OCRScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
$ npx skills add JaidedAI/EasyOCRScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
pix2tex: Using a ViT to convert images of equations into LaTeX code.
$ npx skills add lukas-blecher/LaTeX-OCRScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
GLM-OCR: Accurate × Fast × Comprehensive
$ npx skills add zai-org/GLM-OCRScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
3D Computer Vision Framework
$ npx skills add alicevision/AliceVisionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
An open source library and framework for deep learning on satellite and aerial imagery.
$ npx skills add azavea/raster-visionScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
CLI + Codex · 4 targets
Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs
$ npx skills add Dicklesworthstone/llm_aided_ocrScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
基于PaddleOCR重构,并且脱离PaddlePaddle深度学习训练框架的轻量级OCR,推理速度超快 —— A lightweight OCR system based on PaddleOCR, decoupled from the PaddlePaddle deep learning training framework, wi…
$ npx skills add jingsongliujing/OnnxOCRScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
A Python wrapper for the tesseract-ocr API
$ npx skills add sirfz/tesserocrScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
Rust multi‑backend OCR/VLM engine (DeepSeek‑OCR-1/2, PaddleOCR‑VL, DotsOCR) with DSQ quantization and an OpenAI‑compatible server & CLI – run locally without Python.
$ npx skills add TimmyOVO/deepseek-ocr.rsScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
OpenAI Agents + CLI · 4 targets
zacharywhitley/awesome-ocr is a high-star GitHub project relevant to AI agent workflows.
$ npx skills add zacharywhitley/awesome-ocrScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
We write your reusable computer vision tools. 💜
$ npx skills add roboflow/supervisionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Page 1
Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.
Try the agent resolve APIAI Agent Skill Repository
Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.
Live registry search
Exact name and slug matches are checked against the live registry before ranked alternatives.
Decision filters
Showing 1-16 of 120 ranked candidates matching "macos-vision-ocr"
Best blend of relevance, quality, freshness, and verified outcomes
An iOS OCR Server Using Apple’s Vision Framework
$ npx skills add riddleling/iOS-OCR-ServerScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
Rust library and CLI tool for OCR (extracting text from images)
$ npx skills add robertknight/ocrsScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
让纯文本模型更好地做视觉任务的DeepSeek Harness插件:带意图的图片问答、长截图 OCR、UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, g…
$ npx skills add Anionex/dsh-vision-toolkitScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
Claude Code + CLI · 4 targets
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ lan…
$ npx skills add PaddlePaddle/PaddleOCRScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
OCR software, free and offline. 开源、免费的离线OCR软件。支持截屏/批量导入图片,PDF文档识别,排除水印/页眉页脚,扫描/生成二维码。内置多国语言库。
$ npx skills add hiroi-sora/Umi-OCRScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
$ npx skills add JaidedAI/EasyOCRScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
pix2tex: Using a ViT to convert images of equations into LaTeX code.
$ npx skills add lukas-blecher/LaTeX-OCRScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
GLM-OCR: Accurate × Fast × Comprehensive
$ npx skills add zai-org/GLM-OCRScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
3D Computer Vision Framework
$ npx skills add alicevision/AliceVisionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
An open source library and framework for deep learning on satellite and aerial imagery.
$ npx skills add azavea/raster-visionScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
CLI + Codex · 4 targets
Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs
$ npx skills add Dicklesworthstone/llm_aided_ocrScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
基于PaddleOCR重构,并且脱离PaddlePaddle深度学习训练框架的轻量级OCR,推理速度超快 —— A lightweight OCR system based on PaddleOCR, decoupled from the PaddlePaddle deep learning training framework, wi…
$ npx skills add jingsongliujing/OnnxOCRScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
A Python wrapper for the tesseract-ocr API
$ npx skills add sirfz/tesserocrScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
Rust multi‑backend OCR/VLM engine (DeepSeek‑OCR-1/2, PaddleOCR‑VL, DotsOCR) with DSQ quantization and an OpenAI‑compatible server & CLI – run locally without Python.
$ npx skills add TimmyOVO/deepseek-ocr.rsScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
OpenAI Agents + CLI · 4 targets
zacharywhitley/awesome-ocr is a high-star GitHub project relevant to AI agent workflows.
$ npx skills add zacharywhitley/awesome-ocrScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
We write your reusable computer vision tools. 💜
$ npx skills add roboflow/supervisionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Page 1
Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.
Try the agent resolve API