PaddleSpeech
PaddlePaddle
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System,…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
65–80 / 108
Results: 108
PaddlePaddle
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System,…
supertone-inc
Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
SYSTRAN
Faster Whisper transcription with CTranslate2
abus-aikorea
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTu…
speechbrain
leandromoreira
FFmpeg libav tutorial - learn how media works from basic to transmuxing, transcoding and more. Translations: 🇺🇸 🇨🇳 🇰🇷 🇪🇸 🇻🇳 🇧🇷 🇷🇺
duixcom
🚀 Truly open-source AI avatar(digital human) toolkit for offline video generation and digital human cloning.
espnet
End-to-End Speech Processing Toolkit
gyroflow
Video stabilization using gyroscope data
camenduru
stable diffusion webui colab
kaldi-asr
kaldi-asr/kaldi is the official location of the Kaldi project.
NVlabs
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
debpalash
The open-source ElevenLabs alternative for local voice cloning, design, create, dubbing and dictation Desktop App
Blaizzy
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
open-mmlab
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get…
argmaxinc
On-device Speech AI for Apple Silicon