Text2Video Zero
Picsart-AI-Research
[ICCV 2023 Oral] Text-to-Image Diffusion Models are Zero-Shot Video Generators
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
Page 4 · 16 shown · 223 public entries
Results: 223
Picsart-AI-Research
[ICCV 2023 Oral] Text-to-Image Diffusion Models are Zero-Shot Video Generators
huggingface
Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.
kenshohara
3D ResNets for Action Recognition (CVPR 2018)
justmarkham
Jupyter notebooks from the scikit-learn video series
SandAI-org
MAGI-1: Autoregressive Video Generation at Scale
audeering
Python package for openSMILE
Stonewuu
【融光】 - 基于 Agent 的全流程AI短剧/漫剧/视频创作平台
jy0205
[ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling
bytedance
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
chenyme
这是一个全自动(音频)视频翻译项目。利用Whisper识别声音,AI大模型翻译字幕,最后合并字幕视频,生成翻译后的视频。
Doubiiu
[ECCV 2024, Oral] DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors
williamyang1991
[SIGGRAPH Asia 2023] Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation
enhuiz
An unofficial PyTorch implementation of the audio LM VALL-E
EvolvingLMMs-Lab
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
DamRsn
Audio Plugin for Audio to MIDI transcription using deep learning.
scikit-video