Fun ASR
FunAudioLLM
End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–10 / 10
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 10
FunAudioLLM
End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.
k2-fsa
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Su…
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
espeak-ng
eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.
RHVoice
a free and open source speech synthesizer for Russian and other languages
TensorSpeech
:stuck_out_tongue_closed_eyes: TensorFlowTTS: Real-Time State-of-the-art Speech Synthesis for Tensorflow 2 (supported including English, French, Korean, Chinese, German…
DigitalPhonetics
Controllable and fast Text-to-Speech for over 7000 languages!
Azure-Samples
Microsoft Text-to-Speech API sample code in several languages, part of Cognitive Services.
pnlpal
📚 A customizable dictionary extension that supports double-click lookups in 20+ languages, 1000+ dictionaries, text-to-speech, translation and Anki integration.
VolcanicArts
A modular node-programming language, program creator, animation system, toolkit, router, and debugger made for VRChat