3D Computer Vision Framework
$ npx skills add alicevision/AliceVisionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
AI Agent Skill Repository
Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.
Live registry search
Exact name and slug matches are checked against the live registry before ranked alternatives.
Decision filters
Showing 1-16 of 127 ranked candidates matching "vision-language"
Best blend of relevance, quality, freshness, and verified outcomes
3D Computer Vision Framework
$ npx skills add alicevision/AliceVisionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
An open source library and framework for deep learning on satellite and aerial imagery.
$ npx skills add azavea/raster-visionScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
CLI + Codex · 4 targets
We write your reusable computer vision tools. 💜
$ npx skills add roboflow/supervisionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Datasets, Transforms and Models specific to Computer Vision
$ npx skills add pytorch/visionScenario GitHub automation · I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
CLI + Codex · 4 targets
Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.
$ npx skills add GetStream/Vision-AgentsScenario Coding agents · I need a coding agent that can understand a repository, edit code, and review pull requests.
CLI + Codex · 4 targets
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
$ npx skills add AIDC-AI/OvisScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Practical course about Large Language Models.
$ npx skills add peremartra/Large-Language-Model-Notebooks-CourseScenario RAG and knowledge · I need my agent to build a RAG workflow over documents and retrieve reliable context.
LangChain + CLI · 4 targets
Shared repository for open-sourced projects from the Google AI Language team.
$ npx skills add google-research/languageScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
CLI + Codex · 4 targets
让纯文本模型更好地做视觉任务的DeepSeek Harness插件:带意图的图片问答、长截图 OCR、UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, g…
$ npx skills add Anionex/dsh-vision-toolkitScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
Claude Code + CLI · 4 targets
Node-based Visual Programming Toolbox
$ npx skills add alicevision/MeshroomScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
给纯文本 LLM agent 装上眼睛:图片问答、OCR、截图分析、视觉定位等一套视觉工具箱 + skill,并可无缝接入 Codex、Claude Code、OpenCode、Pi | Give text-only LLM agents vision: image Q&A, OCR, screenshot understanding,…
$ npx skills add Anionex/agent-vision-toolkitScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
Claude Code + OpenAI Agents · 4 targets
Scenic: A Jax Library for Computer Vision Research and Beyond
$ npx skills add google-research/scenicScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
CLI + Codex · 4 targets
The Compute Library is a set of computer vision and machine learning functions optimised for both Arm CPUs and GPUs using SIMD technologies.
$ npx skills add ARM-software/ComputeLibraryScenario Design and creative · I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
CLI + Codex · 4 targets
Computer vision assisted tool to extract numerical data from plot images.
$ npx skills add automeris-io/WebPlotDigitizerScenario Data analysis · I need my agent to analyze CSV data, produce insights, and explain trends.
CLI + Codex · 4 targets
CV-CUDA™ is an open-source, GPU accelerated library for cloud-scale image processing and computer vision.
$ npx skills add CVCUDA/CV-CUDAScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
CLI + Codex · 4 targets
Turn any computer or edge device into a command center for your computer vision projects.
$ npx skills add roboflow/inferenceScenario Coding agents · I need a coding agent that can understand a repository, edit code, and review pull requests.
CLI + Codex · 4 targets
Page 1
Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.
Try the agent resolve APIAI Agent Skill Repository
Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.
Live registry search
Exact name and slug matches are checked against the live registry before ranked alternatives.
Decision filters
Showing 1-16 of 127 ranked candidates matching "vision-language"
Best blend of relevance, quality, freshness, and verified outcomes
3D Computer Vision Framework
$ npx skills add alicevision/AliceVisionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
An open source library and framework for deep learning on satellite and aerial imagery.
$ npx skills add azavea/raster-visionScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
CLI + Codex · 4 targets
We write your reusable computer vision tools. 💜
$ npx skills add roboflow/supervisionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Datasets, Transforms and Models specific to Computer Vision
$ npx skills add pytorch/visionScenario GitHub automation · I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
CLI + Codex · 4 targets
Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.
$ npx skills add GetStream/Vision-AgentsScenario Coding agents · I need a coding agent that can understand a repository, edit code, and review pull requests.
CLI + Codex · 4 targets
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
$ npx skills add AIDC-AI/OvisScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Practical course about Large Language Models.
$ npx skills add peremartra/Large-Language-Model-Notebooks-CourseScenario RAG and knowledge · I need my agent to build a RAG workflow over documents and retrieve reliable context.
LangChain + CLI · 4 targets
Shared repository for open-sourced projects from the Google AI Language team.
$ npx skills add google-research/languageScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
CLI + Codex · 4 targets
让纯文本模型更好地做视觉任务的DeepSeek Harness插件:带意图的图片问答、长截图 OCR、UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, g…
$ npx skills add Anionex/dsh-vision-toolkitScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
Claude Code + CLI · 4 targets
Node-based Visual Programming Toolbox
$ npx skills add alicevision/MeshroomScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
给纯文本 LLM agent 装上眼睛:图片问答、OCR、截图分析、视觉定位等一套视觉工具箱 + skill,并可无缝接入 Codex、Claude Code、OpenCode、Pi | Give text-only LLM agents vision: image Q&A, OCR, screenshot understanding,…
$ npx skills add Anionex/agent-vision-toolkitScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
Claude Code + OpenAI Agents · 4 targets
Scenic: A Jax Library for Computer Vision Research and Beyond
$ npx skills add google-research/scenicScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
CLI + Codex · 4 targets
The Compute Library is a set of computer vision and machine learning functions optimised for both Arm CPUs and GPUs using SIMD technologies.
$ npx skills add ARM-software/ComputeLibraryScenario Design and creative · I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
CLI + Codex · 4 targets
Computer vision assisted tool to extract numerical data from plot images.
$ npx skills add automeris-io/WebPlotDigitizerScenario Data analysis · I need my agent to analyze CSV data, produce insights, and explain trends.
CLI + Codex · 4 targets
CV-CUDA™ is an open-source, GPU accelerated library for cloud-scale image processing and computer vision.
$ npx skills add CVCUDA/CV-CUDAScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
CLI + Codex · 4 targets
Turn any computer or edge device into a command center for your computer vision projects.
$ npx skills add roboflow/inferenceScenario Coding agents · I need a coding agent that can understand a repository, edit code, and review pull requests.
CLI + Codex · 4 targets
Page 1
Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.
Try the agent resolve API