3D Computer Vision Framework
$ npx skills add alicevision/AliceVisionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
AI Agent Skill Repository
Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.
Live registry search
Exact name and slug matches are checked against the live registry before ranked alternatives.
Decision filters
Showing 1-16 of 128 ranked candidates matching "finetuning-vision-models"
Best blend of relevance, quality, freshness, and verified outcomes
3D Computer Vision Framework
$ npx skills add alicevision/AliceVisionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
An open source library and framework for deep learning on satellite and aerial imagery.
$ npx skills add azavea/raster-visionScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
CLI + Codex · 4 targets
Datasets, Transforms and Models specific to Computer Vision
$ npx skills add pytorch/visionScenario GitHub automation · I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
CLI + Codex · 4 targets
We write your reusable computer vision tools. 💜
$ npx skills add roboflow/supervisionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.
$ npx skills add GetStream/Vision-AgentsScenario Coding agents · I need a coding agent that can understand a repository, edit code, and review pull requests.
CLI + Codex · 4 targets
TorchGeo: datasets, samplers, transforms, and pre-trained models for geospatial data
$ npx skills add torchgeo/torchgeoScenario Coding agents · I need a coding agent that can understand a repository, edit code, and review pull requests.
CLI + Codex · 4 targets
让纯文本模型更好地做视觉任务的DeepSeek Harness插件:带意图的图片问答、长截图 OCR、UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, g…
$ npx skills add Anionex/dsh-vision-toolkitScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
Claude Code + CLI · 4 targets
FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit…
$ npx skills add x1xhlol/system-prompts-and-models-of-ai-toolsScenario Coding agents · I need a coding agent that can understand a repository, edit code, and review pull requests.
Claude Code + Cursor · 4 targets
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
$ npx skills add ARahim3/mlx-tuneScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Silero Models: pre-trained text-to-speech models made embarrassingly simple
$ npx skills add snakers4/silero-modelsScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
$ npx skills add open-compass/VLMEvalKitScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
Claude Code + OpenAI Agents · 4 targets
The collection of pre-trained, state-of-the-art AI models for ailia SDK
$ npx skills add ailia-ai/ailia-modelsScenario RAG and knowledge · I need my agent to build a RAG workflow over documents and retrieve reliable context.
CLI + Codex · 4 targets
Node-based Visual Programming Toolbox
$ npx skills add alicevision/MeshroomScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Converted CoreML Model Zoo.
$ npx skills add john-rocky/CoreML-ModelsScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
A repository for storing models that have been inter-converted between various frameworks. Supported frameworks are TensorFlow, PyTorch, ONNX, OpenVINO, TFJS, TFTRT, Ten…
$ npx skills add PINTO0309/PINTO_model_zooScenario Coding agents · I need a coding agent that can understand a repository, edit code, and review pull requests.
CLI + Codex · 4 targets
给纯文本 LLM agent 装上眼睛:图片问答、OCR、截图分析、视觉定位等一套视觉工具箱 + skill,并可无缝接入 Codex、Claude Code、OpenCode、Pi | Give text-only LLM agents vision: image Q&A, OCR, screenshot understanding,…
$ npx skills add Anionex/agent-vision-toolkitScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
Claude Code + OpenAI Agents · 4 targets
Page 1
Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.
Try the agent resolve APIAI Agent Skill Repository
Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.
Live registry search
Exact name and slug matches are checked against the live registry before ranked alternatives.
Decision filters
Showing 1-16 of 128 ranked candidates matching "finetuning-vision-models"
Best blend of relevance, quality, freshness, and verified outcomes
3D Computer Vision Framework
$ npx skills add alicevision/AliceVisionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
An open source library and framework for deep learning on satellite and aerial imagery.
$ npx skills add azavea/raster-visionScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
CLI + Codex · 4 targets
Datasets, Transforms and Models specific to Computer Vision
$ npx skills add pytorch/visionScenario GitHub automation · I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
CLI + Codex · 4 targets
We write your reusable computer vision tools. 💜
$ npx skills add roboflow/supervisionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.
$ npx skills add GetStream/Vision-AgentsScenario Coding agents · I need a coding agent that can understand a repository, edit code, and review pull requests.
CLI + Codex · 4 targets
TorchGeo: datasets, samplers, transforms, and pre-trained models for geospatial data
$ npx skills add torchgeo/torchgeoScenario Coding agents · I need a coding agent that can understand a repository, edit code, and review pull requests.
CLI + Codex · 4 targets
让纯文本模型更好地做视觉任务的DeepSeek Harness插件:带意图的图片问答、长截图 OCR、UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, g…
$ npx skills add Anionex/dsh-vision-toolkitScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
Claude Code + CLI · 4 targets
FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit…
$ npx skills add x1xhlol/system-prompts-and-models-of-ai-toolsScenario Coding agents · I need a coding agent that can understand a repository, edit code, and review pull requests.
Claude Code + Cursor · 4 targets
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
$ npx skills add ARahim3/mlx-tuneScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Silero Models: pre-trained text-to-speech models made embarrassingly simple
$ npx skills add snakers4/silero-modelsScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
$ npx skills add open-compass/VLMEvalKitScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
Claude Code + OpenAI Agents · 4 targets
The collection of pre-trained, state-of-the-art AI models for ailia SDK
$ npx skills add ailia-ai/ailia-modelsScenario RAG and knowledge · I need my agent to build a RAG workflow over documents and retrieve reliable context.
CLI + Codex · 4 targets
Node-based Visual Programming Toolbox
$ npx skills add alicevision/MeshroomScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Converted CoreML Model Zoo.
$ npx skills add john-rocky/CoreML-ModelsScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
A repository for storing models that have been inter-converted between various frameworks. Supported frameworks are TensorFlow, PyTorch, ONNX, OpenVINO, TFJS, TFTRT, Ten…
$ npx skills add PINTO0309/PINTO_model_zooScenario Coding agents · I need a coding agent that can understand a repository, edit code, and review pull requests.
CLI + Codex · 4 targets
给纯文本 LLM agent 装上眼睛:图片问答、OCR、截图分析、视觉定位等一套视觉工具箱 + skill,并可无缝接入 Codex、Claude Code、OpenCode、Pi | Give text-only LLM agents vision: image Q&A, OCR, screenshot understanding,…
$ npx skills add Anionex/agent-vision-toolkitScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
Claude Code + OpenAI Agents · 4 targets
Page 1
Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.
Try the agent resolve API