Dsh Vision Toolkit
Anionex
让纯文本模型更好地做视觉任务的DeepSeek Harness插件:带意图的图片问答、长截图 OCR、UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, g…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
81–96 / 157
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 157
Anionex
让纯文本模型更好地做视觉任务的DeepSeek Harness插件:带意图的图片问答、长截图 OCR、UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, g…
xtreme1-io
Xtreme1 is an all-in-one data labeling and annotation platform for multimodal data training and supports 3D LiDAR point cloud, image, and LLM.
lessthanoptimal
Fast computer vision library for SFM, calibration, fiducials, tracking, image processing, and more.
yanliudesign
Generate original one-ink or controlled two-ink editorial images from any theme, sentence, article idea, object, or reference photo. Always use this skill when the user…
RVC-Boss
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
babysor
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
alphacep
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
PaddlePaddle
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ lan…
openai
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
yzhao062
A Python library for anomaly detection across tabular, time series, graph, text, and image data. 60+ detectors, benchmark-backed ADEngine orchestration, and an agentic w…
enricoros
AI suite powered by state-of-the-art models and providing advanced AI/AGI functions. Includes AI personas, AGI functions, world-class Beam multi-model chats, text-to-ima…
bytedance
The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
mylxsw
An APP that integrates mainstream large language models and image generation models, built with Flutter, with fully open-source code.
crazy-max
Receive notifications when an image is updated on a Docker registry
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.