PaddleSpeech
PaddlePaddle
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System,…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 319
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 319
PaddlePaddle
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System,…
leandromoreira
FFmpeg libav tutorial - learn how media works from basic to transmuxing, transcoding and more. Translations: 🇺🇸 🇨🇳 🇰🇷 🇪🇸 🇻🇳 🇧🇷 🇷🇺
ErickWendel
JS Expert Week 8.0 - 🎥Pre processing videos before uploading in the browser 😏
jik876
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
nyanmisaka
FFmpeg with async and zero-copy Rockchip MPP & RGA support
kan-bayashi
Unofficial Parallel WaveGAN (+ MelGAN & Multi-band MelGAN & HiFi-GAN & StyleMelGAN) with Pytorch
jishengpeng
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
RVC-Boss
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
babysor
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
FoundationVision
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation…
jaywalnut310
VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
unslothai
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.