Lance
bytedance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–12 / 12
Results: 12
bytedance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
EvolvingLMMs-Lab
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
VectorSpaceLab
OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
hao-ai-lab
A unified inference and post-training framework for accelerated video generation.
thu-ml
Official code of Motus: A Unified Latent Action World Model
microsoft
Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
ali-vilab
Code for SCIS-2025 Paper "UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation".
bytedance
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
antgroup
[AAAI 2026] EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation
NVlabs
rCM & Causal-rCM: Leading and Unified Algorithms/Infrastructures for Bidirectional/Autoregressive Video Diffusion Distillation at Scale
X-GenGroup
A unified framework for easy reinforcement learning in Flow-Matching models
HorizonWind2004
[ICLR 2026] Official repo of paper "Reconstruction Alignment Improves Unified Multimodal Models". Unlocking the Massive Zero-shot Potential in Unified Multimodal Models…
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.