1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
每个推荐都保留与其仓库、审计和安装路径的明确关联。
搜索结果: singing-voice-synthesis
英文目录VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
💬 Open source machine learning framework to automate text- and voice-based conversations: NLU, dialogue management, connect to Slack, Facebook, and more - Create chatbots and voice assistants
This skill encodes Emil Kowalski's philosophy on UI polish, component design, animation decisions, and the invisible details that make software feel great.
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
Turn one topic into a finished Vox-style paper-collage explainer/ad video — automated end to end on Atlas Cloud + ffmpeg. An agent skill.
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.
The open-source ElevenLabs alternative for local voice cloning, design, create, dubbing and dictation Desktop App