Agent Vision Toolkit
Anionex
给纯文本 LLM agent 装上眼睛:图片问答、OCR、截图分析、视觉定位等一套视觉工具箱 + skill,并可无缝接入 Codex、Claude Code、OpenCode、Pi | Give text-only LLM agents vision: image Q&A, OCR, screenshot understanding,…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
17–32 / 107
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 107
Anionex
给纯文本 LLM agent 装上眼睛:图片问答、OCR、截图分析、视觉定位等一套视觉工具箱 + skill,并可无缝接入 Codex、Claude Code、OpenCode、Pi | Give text-only LLM agents vision: image Q&A, OCR, screenshot understanding,…
ARM-software
The Compute Library is a set of computer vision and machine learning functions optimised for both Arm CPUs and GPUs using SIMD technologies.
mattpocock
Pressure-test a plan against the codebase, domain glossary, and architectural decisions.
roboflow
A collection of tutorials on state-of-the-art computer vision models and techniques. Explore everything from foundational architectures like ResNet to cutting-edge model…
py-why
DoWhy is a Python library for causal inference that supports explicit modeling and testing of causal assumptions. DoWhy is based on a unified language for causal inferen…
google-ai-edge
Cross-platform, customizable ML solutions for live and streaming media.
roboflow
RF-DETR is a real-time object detection and segmentation model architecture developed by Roboflow, SOTA on COCO, designed for fine-tuning. [ICLR 2026]
emilkowalski
Reverse-lookup glossary that turns a vague description of a web animation or motion effect into its exact term ("the bouncy thing when a popover opens" → Pop in; "the iO…
iamgio
🪐 Markdown with superpowers: from ideas to papers, presentations, websites, books, and knowledge bases.
Developer-Y
List of Computer Science courses with video lectures.
argotorg
Solidity, the Smart Contract Programming Language
Tencent-Hunyuan
HunyuanVideo: A Systematic Framework For Large Video Generation Model