UniAnimate
ali-vilab
Code for SCIS-2025 Paper "UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation".
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
97–112 / 414
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 414
ali-vilab
Code for SCIS-2025 Paper "UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation".
cure-lab
[ICLR24] Official implementation of the paper “MagicDrive: Street View Generation with Diverse 3D Geometry Control”
ermongroup
PyTorch implementation for SDEdit: Image Synthesis and Editing with Stochastic Differential Equations
shadow2496
Official PyTorch implementation of "VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization" (CVPR 2021)
zai-org
CogView4, CogView3-Plus and CogView3(ECCV 2024)
alumae
Real-time full-duplex speech recognition server, based on the Kaldi toolkit and the GStreamer framwork.
videoflow
Python framework that facilitates the quick development of complex video analysis applications and other series-processing based applications in a multiprocessing enviro…
Anil-matcha
🎬 Turn any topic into a finished Vox-style paper-collage explainer / motion graphics video — script, collage keyframes, animation, voice-over, music & captions, all aut…
Aratako
A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control
alesaccoia
Near-Realtime audio transcription using self-hosted Whisper and WebSocket in Python/JS
Vonage
Vonage REST API client for PHP. API support for SMS, Voice, Text-to-Speech, Numbers, Verify (2FA) and more.
bytedance
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
EvolvingLMMs-Lab
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.