GPT SoVITS
RVC-Boss
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
Page 1 · 16 shown · 629 public entries
Results: 629
RVC-Boss
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
google-ai-edge
Cross-platform, customizable ML solutions for live and streaming media.
unslothai
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
ggml-org
Port of OpenAI's Whisper model in C/C++
alphacep
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
huggingface
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
OpenBMB
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
FunAudioLLM
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
index-tts
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
jianchang512
Translate the video from one language to another and embed dubbing & subtitles.
babysor
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
k2-fsa
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Su…
PaddlePaddle
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System,…
supertone-inc
Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.