Nano World Model
simchowitzlabpublic
A Minimalist, Batteries-included Repository for Advancing World Model Science.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 57
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 57
simchowitzlabpublic
A Minimalist, Batteries-included Repository for Advancing World Model Science.
ModelTC
Lightweight Image Video Action Generation Inference Framework
Tencent-Hunyuan
HunyuanVideo: A Systematic Framework For Large Video Generation Model
vllm-project
A framework for efficient model inference with omni-modality models
FoundationVision
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation…
PaddlePaddle
PaddlePaddle GAN library, including lots of interesting applications like First-Order motion transfer, Wav2Lip, picture repair, image editing, photo2cartoon, image style…
open-mmlab
OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion mod…
thu-ml
TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
PKU-YuanGroup
Helios: Real Real-Time Long Video Generation Model
MoonInTheRiver
DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (SVS & TTS); AAAI 2022; Official code
Tencent-Hunyuan
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
JIA-Lab-research
This project is the official implementation of 'DreamOmni2: Multimodal Instruction-based Editing and Generation (CVPR2026 Highlight)''
FunAudioLLM
End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.
shivammehta25
[ICASSP 2024] 🍵 Matcha-TTS: A fast TTS architecture with conditional flow matching