OmniGen
VectorSpaceLab
OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 20
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 20
VectorSpaceLab
OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
MahmoudAshraf97
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
sanchit-gandhi
JAX implementation of OpenAI's Whisper model for up to 70x speed-up on TPU.
williamyang1991
[SIGGRAPH Asia 2023] Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation
harskish
Discovering Interpretable GAN Controls [NeurIPS 2020]
snap-research
Code for Motion Representations for Articulated Animation paper
alphacep
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
camenduru
stable diffusion webui colab
bloc97
A High-Quality Real Time Upscaler for Anime Video
snakers4
Silero Models: pre-trained text-to-speech models made embarrassingly simple
NVIDIA
State-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible accuracy and performance on enterprise-grade infrastructure.
FoundationVision
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation…
mozilla
:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)
open-mmlab
OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion mod…
shivammehta25
[ICASSP 2024] 🍵 Matcha-TTS: A fast TTS architecture with conditional flow matching