Amphion
open-mmlab
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
49–64 / 365
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 365
open-mmlab
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get…
snap-research
Code for Motion Representations for Articulated Animation paper
leejet
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
snakers4
Silero Models: pre-trained text-to-speech models made embarrassingly simple
nl8590687
A Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统
NVIDIA
State-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible accuracy and performance on enterprise-grade infrastructure.
modelscope
Open-source, accurate and easy-to-use video speech recognition & clipping tool. LLM-based AI clipping integrated.
junyanz
Software that can generate photos from paintings, turn horses into zebras, perform style transfer, and more.
wenet-e2e
Production First and Production Ready End-to-End Speech Recognition Toolkit
vllm-project
A framework for efficient model inference with omni-modality models
remsky
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model w/multiplatform CPU, AMD, NVIDIA GPU PyTorch support, handling, and auto-stitching
Breakthrough
:movie_camera: Python and OpenCV-based scene cut/transition detection program & library.
Picovoice
On-device wake word detection powered by deep learning
espeak-ng
eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.