OpenVoice
myshell-ai
Instant voice cloning by MIT and MyShell. Audio foundation model.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
17–32 / 143
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 143
myshell-ai
Instant voice cloning by MIT and MyShell. Audio foundation model.
rany2
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
Uberi
Speech recognition module for Python, supporting several engines and APIs, online and offline.
Blaizzy
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
zai-org
text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
open-mmlab
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get…
snakers4
Silero Models: pre-trained text-to-speech models made embarrassingly simple
remsky
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model w/multiplatform CPU, AMD, NVIDIA GPU PyTorch support, handling, and auto-stitching
denizsafak
Generate audiobooks from EPUBs, PDFs and text with synchronized captions.
espeak-ng
eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.
mozilla
:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)
MahmoudAshraf97
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper