HunyuanImage 3.0
Tencent-Hunyuan
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
17β32 / 60
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 60
Tencent-Hunyuan
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
JIA-Lab-research
This project is the official implementation of 'DreamOmni2: Multimodal Instruction-based Editing and Generation (CVPR2026 Highlight)''
FunAudioLLM
End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.
shivammehta25
[ICASSP 2024] π΅ Matcha-TTS: A fast TTS architecture with conditional flow matching
bytedance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
FoundationVision
[CVPR 2025 Oral]Infinity β : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
OpenGVLab
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, Sβ¦
FireRedTeam
FireRed-Image-Edit is a powerful image editing foundation model achieving open-source state-of-the-art performance with precise instruction following, high-fidelity geneβ¦
kandinskylab
Kandinsky 5.0: A family of diffusion models for Video & Image generation
FoundationVision
Autoregressive Model Beats Diffusion: π¦ Llama for Scalable Image Generation
thu-ml
Official code of Motus: A Unified Latent Action World Model
ali-vilab
Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
snap-research
Code for Motion Representations for Articulated Animation paper
Aratako
A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control