CosyVoice
FunAudioLLM
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–6 / 6
Results: 6
FunAudioLLM
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
gnipbao
Agent skill: convert Chinese story copy or ordered images into a hand-drawn diary-comic animation (silent MP4 picture track).
k2-fsa
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Su…
open-mmlab
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get…
remsky
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model w/multiplatform CPU, AMD, NVIDIA GPU PyTorch support, handling, and auto-stitching
Pluviobyte
Render, preview, validate, and burn consistent production subtitles from quality-checked timestamp artifacts.