Voice Pro
abus-aikorea
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTu…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
113–128 / 278
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 278
abus-aikorea
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTu…
junyanz
Image-to-Image Translation in PyTorch
gradio-app
The python library for real-time communication
alesaccoia
Near-Realtime audio transcription using self-hosted Whisper and WebSocket in Python/JS
NVlabs
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
debpalash
The open-source ElevenLabs alternative for local voice cloning, design, create, dubbing and dictation Desktop App
Blaizzy
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
open-mmlab
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get…
modelscope
Open-source, accurate and easy-to-use video speech recognition & clipping tool. LLM-based AI clipping integrated.
BinWang28
The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
wenet-e2e
Production First and Production Ready End-to-End Speech Recognition Toolkit
vllm-project
A framework for efficient model inference with omni-modality models
remsky
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model w/multiplatform CPU, AMD, NVIDIA GPU PyTorch support, handling, and auto-stitching
denizsafak
Generate audiobooks from EPUBs, PDFs and text with synchronized captions.