Pdf Craft
oomol-lab
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
33–48 / 200
Results: 200
oomol-lab
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
ramjke
Advanced real-time screen translator for games, hardcoded subtitles in videos, static text and etc.
DayBreak-u
超轻量级中文ocr,支持竖排文字识别, 支持ncnn、mnn、tnn推理 ( dbnet(1.8M) + crnn(2.5M) + anglenet(378KB)) 总模型仅4.7M
hiroi-sora
OCR software, free and offline. 开源、免费的离线OCR软件。支持截屏/批量导入图片,PDF文档识别,排除水印/页眉页脚,扫描/生成二维码。内置多国语言库。
A9T9
Ui.Vision Open-Source RPA Software with Computer Vision, OCR, Anthropic Computer Use/LLM. Selenium IDE import/export.
grobidOrg
A machine learning software for extracting information from scholarly documents
datalab-to
OCR model that handles complex tables, forms, handwriting with full layout.
dmMaze
深度学习辅助漫画翻译工具, 支持一键机翻和简单的图像/文本编辑 | Yet another computer-aided comic/manga translation tool powered by deeplearning
TheJoeFin
Use OCR in Windows quickly and easily with Text Grab. With optional background process and notifications.
KnpLabs
PHP library allowing thumbnail, snapshot or PDF generation from a url or a html page. Wrapper for wkhtmltopdf/wkhtmltoimage
JabRef
Graphical Java application for managing BibTeX and BibLaTeX (.bib) databases
ruvnet
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
umlx5h
The media player for language learning, with dual subtitles, AI-generated subtitles, real-time translation, and more!
Anionex
让纯文本模型更好地做视觉任务的DeepSeek Harness插件:带意图的图片问答、长截图 OCR、UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, g…
JaidedAI
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
deepdoctection
Design direction, Figma implementation, UI review, React performance, browser QA, and safe preview deployments.
Code review, repo inspection, testing, planning, shipping, and engineering workflows for coding agents.
Research, recent-events briefings, PDF parsing, markdown conversion, RAG ingestion, and knowledge workflows.
Crawling, scraping, extraction, browser automation, structured data capture, and website-to-markdown workflows.
Presentation generation, editable PPTX, slide decks, speaker notes, and visual storytelling workflows.
Video-generation prompts, B-roll, Vox-style explainers, camera direction, captions, and creative-production workflows.
Image, video, creative production, UI design, multimodal generation, and visual workflow skills.
Data analysis, analytics, ETL, notebooks, databases, tables, charts, and reporting workflows.