Dia
nari-labs
A TTS model capable of generating ultra-realistic dialogue in one pass.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
49–64 / 77
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 77
nari-labs
A TTS model capable of generating ultra-realistic dialogue in one pass.
myshell-ai
Instant voice cloning by MIT and MyShell. Audio foundation model.
leejet
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
remsky
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model w/multiplatform CPU, AMD, NVIDIA GPU PyTorch support, handling, and auto-stitching
OpenMOSS
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and…
Tencent-Hunyuan
HunyuanVideo-1.5: A leading lightweight video generation model
sanchit-gandhi
JAX implementation of OpenAI's Whisper model for up to 70x speed-up on TPU.
SkyworkAI
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
metavoiceio
Foundational model for human-like, expressive TTS
lenML
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.
ai-forever
Kandinsky 2 — multilingual text2image latent diffusion model
yeyupiaoling
Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inf…
Enemyx-net
A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly within your C…
thu-ml
[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.