OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 22
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 22
vllm-project
A framework for efficient model inference with omni-modality models
fikrikarim
On-device, real-time multimodal AI. Have natural voice and vision conversations with an AI that runs entirely on your machine. Powered by Gemma 4 E2B and Kokoro.
Tencent-Hunyuan
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
bytedance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
EvolvingLMMs-Lab
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
zjukg
Knowledge Graphs Meet Multi-Modal Learning: A Comprehensive Survey
open-mmlab
OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion mod…
bytedance
SALMONN family: A suite of advanced multi-modal LLMs
JIA-Lab-research
This project is the official implementation of 'DreamOmni2: Multimodal Instruction-based Editing and Generation (CVPR2026 Highlight)''
lhotse-speech
Tools for handling multimodal data in machine learning projects.
OpenGVLab
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, S…
Phantom-video
Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
Tencent-Hunyuan
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
microsoft
Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
antgroup
[AAAI 2026] EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation