SpeechT5
microsoft
Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
337–352 / 478
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 478
microsoft
Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
MiteshPuthran
The neural network model is capable of detecting five different male/female emotions from audio speeches. (Deep Learning, NLP, Python)
coqui-ai
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
ali-vilab
Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
receyuki
A simple standalone viewer for reading prompts from Stable Diffusion generated image outside the webui.
CSTR-Edinburgh
This is now the official location of the Merlin project.
mit-han-lab
[NeurIPS 2020] Differentiable Augmentation for Data-Efficient GAN Training
Uminosachi
Inpaint Anything extension performs stable diffusion inpainting on a browser UI using masks from Segment Anything.
ictnlp
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
alphacep
WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries
omerbt
Official Pytorch Implementation for "MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation" presenting "MultiDiffusion" (ICML 2023)
showlab
[ECCV 2024 Oral] MotionDirector: Motion Customization of Text-to-Video Diffusion Models.
TensorSpeech
:zap: TensorFlowASR: Almost State-of-the-art Automatic Speech Recognition in Tensorflow 2. Supported languages that can use characters or subwords
ManimCommunity
Manim plugin for all things voiceover