MagicDrive
cure-lab
[ICLR24] Official implementation of the paper “MagicDrive: Street View Generation with Diverse 3D Geometry Control”
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
97–112 / 271
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 271
cure-lab
[ICLR24] Official implementation of the paper “MagicDrive: Street View Generation with Diverse 3D Geometry Control”
ermongroup
PyTorch implementation for SDEdit: Image Synthesis and Editing with Stochastic Differential Equations
shadow2496
Official PyTorch implementation of "VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization" (CVPR 2021)
zai-org
CogView4, CogView3-Plus and CogView3(ECCV 2024)
alumae
Real-time full-duplex speech recognition server, based on the Kaldi toolkit and the GStreamer framwork.
showlab
[ECCV 2024 Oral] MotionDirector: Motion Customization of Text-to-Video Diffusion Models.
videoflow
Python framework that facilitates the quick development of complex video analysis applications and other series-processing based applications in a multiprocessing enviro…
Aratako
A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control
alesaccoia
Near-Realtime audio transcription using self-hosted Whisper and WebSocket in Python/JS
BinWang28
The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
AutoArk
[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!
bytedance
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
EvolvingLMMs-Lab
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
kandinskylab
Kandinsky 5.0: A family of diffusion models for Video & Image generation
antgroup
[AAAI 2026] EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation