HunyuanImage 3.0
Tencent-Hunyuan
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
33–48 / 91
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 91
Tencent-Hunyuan
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
JIA-Lab-research
This project is the official implementation of 'DreamOmni2: Multimodal Instruction-based Editing and Generation (CVPR2026 Highlight)''
shivammehta25
[ICASSP 2024] 🍵 Matcha-TTS: A fast TTS architecture with conditional flow matching
bytedance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
OpenGVLab
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, S…
FireRedTeam
FireRed-Image-Edit is a powerful image editing foundation model achieving open-source state-of-the-art performance with precise instruction following, high-fidelity gene…
Eyeline-Labs
Official code, models, and data for Vista4D: Video Reshooting with 4D Point Clouds (CVPR 2026 Highlight)
ali-vilab
Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
snap-research
Code for Motion Representations for Articulated Animation paper
omerbt
Official Pytorch Implementation for "MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation" presenting "MultiDiffusion" (ICML 2023)
Aratako
A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control
thu-ml
[ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation" & Causal For…
NVlabs
rCM & Causal-rCM: Leading and Unified Algorithms/Infrastructures for Bidirectional/Autoregressive Video Diffusion Distillation at Scale
Picovoice
On-device Speech-to-Intent engine powered by deep learning
RVC-Boss
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)