Vision Agents
GetStream
Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–13 / 13
Results: 13
GetStream
Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.
snakers4
Silero Models: pre-trained text-to-speech models made embarrassingly simple
ailia-ai
The collection of pre-trained, state-of-the-art AI models for ailia SDK
huggingface
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
Robbyant
Advancing Open-source World Models
unslothai
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
invoke-ai
Invoke is a leading creative engine for Stable Diffusion models, empowering professionals, artists, and enthusiasts to generate and create visual media using the latest…
Wan-Video
Wan: Open and Advanced Large-Scale Video Generative Models
vllm-project
A framework for efficient model inference with omni-modality models
huggingface
Build local voice agents with open-source models
zebbern
80+ free AI services for chat, image, video, voice & APIs (may sometimes include access to lead gen ai models for free)
fikrikarim
On-device, real-time multimodal AI. Have natural voice and vision conversations with an AI that runs entirely on your machine. Powered by Gemma 4 E2B and Kokoro.
NVIDIA-AI-Blueprints
The NVIDIA VSS Blueprint is a suite of reference architectures for building GPU-accelerated vision agents and AI-powered video analytics applications.