Text2Video Zero
Picsart-AI-Research
[ICCV 2023 Oral] Text-to-Image Diffusion Models are Zero-Shot Video Generators
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
17–32 / 335
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 335
Picsart-AI-Research
[ICCV 2023 Oral] Text-to-Image Diffusion Models are Zero-Shot Video Generators
FireRedTeam
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering ou…
nvidia-cosmos
Cosmos-Predict2.5, the latest version of the Cosmos World Foundation Models (WFMs) family, specialized for simulating and predicting the future state of the world in the…
s60sc
ESP32 Camera motion capture application to record JPEGs to SD card as AVI files and stream to browser as MJPEG. If a microphone is installed then a WAV file is also crea…
lhotse-speech
Tools for handling multimodal data in machine learning projects.
OpenGVLab
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, S…
pluja
Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!
diodiogod
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio…
readbeyond
aeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced alignment)
PKU-YuanGroup
[TPAMI 2025🔥] MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators
coqui-ai
🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.
GauravSingh9356
Personal Assistant built using python libraries. It does almost anything which includes sending emails, Optical Text Recognition, Dynamic News Reporting at any time with…
harskish
Discovering Interpretable GAN Controls [NeurIPS 2020]
ali-vilab
Official implementations for paper: DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic Models
lucidrains
Implementation of Video Diffusion Models, Jonathan Ho's new paper extending DDPMs to Video Generation - in Pytorch