Wav2letter
flashlight
Facebook AI Research's Automatic Speech Recognition Toolkit
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
33–48 / 115
Results: 115
flashlight
Facebook AI Research's Automatic Speech Recognition Toolkit
mkiol
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.
huggingface
Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.
FunAudioLLM
End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.
royshil
OBS plugin for local speech recognition and captioning using AI
k2-fsa
Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, Windows,…
yeyupiaoling
Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inf…
mravanelli
pytorch-kaldi is a project for developing state-of-the-art DNN/RNN hybrid speech recognition systems. The DNN part is managed by pytorch, while feature extraction, label…
nobody132
中文语音识别; Mandarin Automatic Speech Recognition;
syhw
Attempt at tracking states of the arts and recent results (bibliography) on speech recognition.
sooftware
[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)
alphacep
Offline speech recognition for Android with Vosk library.
sdkcarlos
A voice control - voice commands - speech recognition and speech synthesis javascript library. Create your own siri,google now or cortana with Google Chrome within your…
alphacep
WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries
babysor
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
k2-fsa
Speech-to-text server framework with next-gen Kaldi