MIMO
menyifang
Official implementation of "MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling"
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
321–336 / 478
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 478
menyifang
Official implementation of "MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling"
Saik0s
The open-source iOS app that's making quality voice transcription more accessible on mobile devices.
microsoft
Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
MiteshPuthran
The neural network model is capable of detecting five different male/female emotions from audio speeches. (Deep Learning, NLP, Python)
coqui-ai
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
ali-vilab
Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
receyuki
A simple standalone viewer for reading prompts from Stable Diffusion generated image outside the webui.
CSTR-Edinburgh
This is now the official location of the Merlin project.
mit-han-lab
[NeurIPS 2020] Differentiable Augmentation for Data-Efficient GAN Training
jishengpeng
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
Uminosachi
Inpaint Anything extension performs stable diffusion inpainting on a browser UI using masks from Segment Anything.
ictnlp
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
alphacep
WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries
gitmylo
A webui for different audio related Neural Networks
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.