IMS Toucan
DigitalPhonetics
Controllable and fast Text-to-Speech for over 7000 languages!
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
65–80 / 161
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 161
DigitalPhonetics
Controllable and fast Text-to-Speech for over 7000 languages!
pluja
Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!
diodiogod
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio…
readbeyond
aeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced alignment)
ai-forever
Kandinsky 2 — multilingual text2image latent diffusion model
R3gm
Synchronized Translation for Videos. Video dubbing
PKU-YuanGroup
[TPAMI 2025🔥] MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators
coqui-ai
🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.
marytts
MARY TTS -- an open-source, multilingual text-to-speech synthesis system written in pure java
jik876
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
Rayhane-mamah
DeepMind's Tacotron-2 Tensorflow implementation
pannous
🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks