SenseVoice
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–9 / 9
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 9
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
Breakthrough
:movie_camera: Python and OpenCV-based scene cut/transition detection program & library.
Picovoice
On-device wake word detection powered by deep learning
AlekPet
Custom nodes that extend the capabilities of Comfyui
sandrohanea
Whisper.net. Speech to text made simple using Whisper Models
FireRedTeam
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switchin…
k2-fsa
Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, Windows,…