SenseVoice
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–10 / 10
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 10
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
Breakthrough
:movie_camera: Python and OpenCV-based scene cut/transition detection program & library.
Picovoice
On-device wake word detection powered by deep learning
AlekPet
Custom nodes that extend the capabilities of Comfyui
sandrohanea
Whisper.net. Speech to text made simple using Whisper Models
ManimCommunity
Manim plugin for all things voiceover
FireRedTeam
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switchin…
k2-fsa
Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, Windows,…