SenseVoice
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
17–32 / 73
Results: 73
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
huggingface
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and traini…
abus-aikorea
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTu…
open-mmlab
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get…
spotify
A lightweight yet powerful audio-to-MIDI converter with pitch bend detection
rsxdalv
A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Au…
milvus-io
Dealing with all unstructured data, such as reverse image search, audio search, molecular search, video analysis, question and answer systems, NLP, etc.
pluja
Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!
enhuiz
An unofficial PyTorch implementation of the audio LM VALL-E
readbeyond
aeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced alignment)
DamRsn
Audio Plugin for Audio to MIDI transcription using deep learning.
MiteshPuthran
The neural network model is capable of detecting five different male/female emotions from audio speeches. (Deep Learning, NLP, Python)
jishengpeng
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
mravanelli
SincNet is a neural architecture for efficiently processing raw audio samples.
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.