Text2Video Zero
Picsart-AI-Research
[ICCV 2023 Oral] Text-to-Image Diffusion Models are Zero-Shot Video Generators
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
49–64 / 244
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 244
Picsart-AI-Research
[ICCV 2023 Oral] Text-to-Image Diffusion Models are Zero-Shot Video Generators
ali-vilab
Official implementations for paper: Anydoor: zero-shot object-level image customization
huggingface
Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.
JIA-Lab-research
This project is the official implementation of 'DreamOmni2: Multimodal Instruction-based Editing and Generation (CVPR2026 Highlight)''
Janspiry
Unofficial implementation of Image Super-Resolution via Iterative Refinement by Pytorch
SandAI-org
MAGI-1: Autoregressive Video Generation at Scale
nvidia-cosmos
Cosmos-Predict2.5, the latest version of the Cosmos World Foundation Models (WFMs) family, specialized for simulating and predicting the future state of the world in the…
Tencent
High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
numz
Official SeedVR2 Video Upscaler for ComfyUI
bytedance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
ThioJoe
Automatically translates the text of a video based on a subtitle file, and then uses AI voice services to create a new dubbed & translated audio track where the speech i…
janvarev
Ирина - русский голосовой ассистент для работы оффлайн. Поддерживает скиллы через плагины.
lhotse-speech
Tools for handling multimodal data in machine learning projects.
jy0205
[ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling