Audio AI Hub
BinWang28
The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 40
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 40
BinWang28
The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
FoundationVision
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation…
ali-vilab
[ICCV 2025] Official implementations for paper: VACE: All-in-One Video Creation and Editing
ali-vilab
Official implementations for paper: Anydoor: zero-shot object-level image customization
k2-fsa
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Su…
openvinotoolkit
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
coqui-ai
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
duixcom
🚀 Truly open-source AI avatar(digital human) toolkit for offline video generation and digital human cloning.
ali-vilab
Official implementations for paper: DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic Models
VectorSpaceLab
OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
debpalash
The open-source ElevenLabs alternative for local voice cloning, design, create, dubbing and dictation Desktop App
lucidrains
Implementation of Video Diffusion Models, Jonathan Ho's new paper extending DDPMs to Video Generation - in Pytorch
open-mmlab
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get…
snap-research
Code for Motion Representations for Articulated Animation paper
modelscope
Open-source, accurate and easy-to-use video speech recognition & clipping tool. LLM-based AI clipping integrated.
espeak-ng
eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.