No visual example yet
Explore the skillDiffusers
huggingface
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–12 / 12
Results: 12
No visual example yet
Explore the skillhuggingface
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
No visual example yet
Explore the skillOpenBMB
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
No visual example yet
Explore the skillFunAudioLLM
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
No visual example yet
Explore the skillindex-tts
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
No visual example yet
Explore the skillk2-fsa
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Su…
No visual example yet
Explore the skillrany2
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
No visual example yet
Explore the skillBlaizzy
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
No visual example yet
Explore the skillsnakers4
Silero Models: pre-trained text-to-speech models made embarrassingly simple
No visual example yet
Explore the skillremsky
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model w/multiplatform CPU, AMD, NVIDIA GPU PyTorch support, handling, and auto-stitching
No visual example yet
Explore the skilldenizsafak
Generate audiobooks from EPUBs, PDFs and text with synchronized captions.
No visual example yet
Explore the skillhuggingface
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and traini…
No visual example yet
Explore the skillPaddlePaddle
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System,…