Open Speech Corpora
coqui-ai
π A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
241β256 / 478
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 478
coqui-ai
π A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
lucidrains
Implementation of Video Diffusion Models, Jonathan Ho's new paper extending DDPMs to Video Generation - in Pytorch
ali-vilab
Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
mayuelala
[AAAI 2024] Follow-Your-Pose: This repo is the official implementation of "Follow-Your-Pose : Pose-Guided Text-to-Video Generation using Pose-Free Videos"
receyuki
A simple standalone viewer for reading prompts from Stable Diffusion generated image outside the webui.
CSTR-Edinburgh
This is now the official location of the Merlin project.
mit-han-lab
[NeurIPS 2020] Differentiable Augmentation for Data-Efficient GAN Training
jishengpeng
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
snap-research
Code for Motion Representations for Articulated Animation paper
ictnlp
StreamSpeech is an βAll in Oneβ seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
huanngzh
[ICCV 2025] Official impl. of "MV-Adapter: Multi-view Consistent Image Generation Made Easy"
sdkcarlos
A voice control - voice commands - speech recognition and speech synthesis javascript library. Create your own siri,google now or cortana with Google Chrome within yourβ¦
alphacep
WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries
gitmylo
A webui for different audio related Neural Networks
mravanelli
SincNet is a neural architecture for efficiently processing raw audio samples.