TTS WebUI
rsxdalv
A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Au…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
17–32 / 55
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 55
rsxdalv
A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Au…
OpenShot
OpenShot Video Library (libopenshot) is a free, open-source project dedicated to delivering high quality video editing, animation, and playback solutions to the world. A…
bytedance
SALMONN family: A suite of advanced multi-modal LLMs
huggingface
Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.
devnen
Self-host the powerful Chatterbox TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), predefined voices, voice clonin…
letoram
Arcan - [Display Server, Multimedia Framework, Game Engine] -> "Desktop Engine"
lhotse-speech
Tools for handling multimodal data in machine learning projects.
pluja
Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!
enhuiz
An unofficial PyTorch implementation of the audio LM VALL-E
readbeyond
aeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced alignment)
zzw922cn
End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow
R3gm
Synchronized Translation for Videos. Video dubbing
Enemyx-net
A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly within your C…
julius-speech
Open-Source Large Vocabulary Continuous Speech Recognition Engine