No visual example yet
Explore the skillStyleTTS2
yl4579
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
9 Skills
Results: 9
No visual example yet
Explore the skillyl4579
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
No visual example yet
Explore the skillhuggingface
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and traini…
No visual example yet
Explore the skillunslothai
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
No visual example yet
Explore the skillFunAudioLLM
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
No visual example yet
Explore the skillNVIDIA
A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference appl…
No visual example yet
Explore the skillhkchengrex
[CVPR 2025] MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
No visual example yet
Explore the skillcoqui-ai
🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.
No visual example yet
Explore the skillyeyupiaoling
Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inf…
No visual example yet
Explore the skillFinrandojin
AI-powered multi-voice audiobook generator — LLM script annotation, voice cloning, voice design, LoRA training, per-line style control, and export to MP3, chaptered M4B,…