Wav2letter
flashlight
Facebook AI Research's Automatic Speech Recognition Toolkit
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
17–32 / 46
Results: 46
flashlight
Facebook AI Research's Automatic Speech Recognition Toolkit
Azure-Samples
Microsoft Text-to-Speech API sample code in several languages, part of Cognitive Services.
jianchang512
Voice Recognition to Text Tool / 一个离线运行的本地音视频转字幕工具,输出json、srt字幕、纯文字格式
mkiol
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.
FireRedTeam
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering ou…
FunAudioLLM
End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.
yeyupiaoling
Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inf…
rany2
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
babysor
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
OpenBMB
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
index-tts
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
k2-fsa
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Su…
coqui-ai
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
open-mmlab
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get…
snakers4
Silero Models: pre-trained text-to-speech models made embarrassingly simple
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.