Normcap
dynobo
OCR powered screen-capture tool to capture information instead of images
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
33–48 / 69
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 69
dynobo
OCR powered screen-capture tool to capture information instead of images
killkimno
MORT 번역기 프로젝트 - Real-time game translator with OCR
otiai10
Go package for OCR (Optical Character Recognition), by using Tesseract C++ library
NanoNets
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
AlibabaResearch
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab,…
ttop32
Mouseover Translate Any Language At Once - Chrome Extension: PDF Translator, EBOOK, EPUB, OCR, TTS, NETFLIX, YOUTUBE DUAL SUBTITLES, GOOGLE DOCS, AI, VIEWER, GMAIL, WRIT…
manisandro
A Gtk/Qt front-end to tesseract-ocr.
NanoNets
Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extract…
opendatalab
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
getomni-ai
OCR & Document Extraction using vision models
Sumanth077
A curated collection of practical AI projects implementing OCR systems, RAG, AI agents, and other AI use cases.
SkyworkAI
DeepResearchAgent is a hierarchical multi-agent system designed not only for deep research tasks but also for general-purpose task solving. The framework leverages a top…
clovaai
Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generator (SynthDoG), ECCV 2022
Anionex
让纯文本模型更好地做视觉任务的DeepSeek Harness插件:带意图的图片问答、长截图 OCR、UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, g…