πClone a voice in 5 seconds to generate arbitrary speech in real-time
$ npx skills add babysor/MockingBirdAlternatives
Compare similar skills by workflow fit, trust score, quality, GitHub adoption, maintenance, and install readiness.
Current skill
An unofficial PyTorch implementation of the audio LM VALL-E
πClone a voice in 5 seconds to generate arbitrary speech in real-time
$ npx skills add babysor/MockingBird1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
$ npx skills add RVC-Boss/GPT-SoVITSVoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
$ npx skills add OpenBMB/VoxCPMTranslate the video from one language to another and embed dubbing & subtitles.
$ npx skills add jianchang512/pyvideotransAn Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
$ npx skills add index-tts/index-ttsWhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
$ npx skills add m-bain/whisperXπ AI ε ¨θͺε¨ηθ§ι’εΌζ | AI Fully Automated Short Video Engine
$ npx skills add AIDC-AI/Pixelle-VideoUnsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
$ npx skills add unslothai/unslothSpeech recognition module for Python, supporting several engines and APIs, online and offline.
$ npx skills add Uberi/speech_recognitionSpeech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages
$ npx skills add k2-fsa/sherpa-onnxA PyTorch-based Speech Toolkit
$ npx skills add speechbrain/speechbrainLightning-Fast, On-Device, Multilingual TTS β running natively via ONNX.
$ npx skills add supertone-inc/supertonicEnd-to-End Speech Processing Toolkit
$ npx skills add espnet/espnetOpenAI Whisper ASR Webservice API
$ npx skills add ahmetoner/whisper-asr-webservicePort of OpenAI's Whisper model in C/C++
$ npx skills add ggml-org/whisper.cppπ€ Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
$ npx skills add huggingface/diffusersHow to choose
Use an alternative when it has a clearer install path, higher trust score, fresher maintenance, or better platform fit for your current agent stack. Keep Vall E if it already passes your workflow test and repository review.
Next step
Open the compare page, test the install commands in a sandbox, and check each repository before using a skill in production.
Alternatives
Compare similar skills by workflow fit, trust score, quality, GitHub adoption, maintenance, and install readiness.
Current skill
An unofficial PyTorch implementation of the audio LM VALL-E
πClone a voice in 5 seconds to generate arbitrary speech in real-time
$ npx skills add babysor/MockingBird1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
$ npx skills add RVC-Boss/GPT-SoVITSVoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
$ npx skills add OpenBMB/VoxCPMTranslate the video from one language to another and embed dubbing & subtitles.
$ npx skills add jianchang512/pyvideotransAn Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
$ npx skills add index-tts/index-ttsWhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
$ npx skills add m-bain/whisperXπ AI ε ¨θͺε¨ηθ§ι’εΌζ | AI Fully Automated Short Video Engine
$ npx skills add AIDC-AI/Pixelle-VideoUnsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
$ npx skills add unslothai/unslothSpeech recognition module for Python, supporting several engines and APIs, online and offline.
$ npx skills add Uberi/speech_recognitionSpeech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages
$ npx skills add k2-fsa/sherpa-onnxA PyTorch-based Speech Toolkit
$ npx skills add speechbrain/speechbrainLightning-Fast, On-Device, Multilingual TTS β running natively via ONNX.
$ npx skills add supertone-inc/supertonicEnd-to-End Speech Processing Toolkit
$ npx skills add espnet/espnetOpenAI Whisper ASR Webservice API
$ npx skills add ahmetoner/whisper-asr-webservicePort of OpenAI's Whisper model in C/C++
$ npx skills add ggml-org/whisper.cppπ€ Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
$ npx skills add huggingface/diffusersHow to choose
Use an alternative when it has a clearer install path, higher trust score, fresher maintenance, or better platform fit for your current agent stack. Keep Vall E if it already passes your workflow test and repository review.
Next step
Open the compare page, test the install commands in a sandbox, and check each repository before using a skill in production.