No visual example yet
Explore the skillTTS Audio Suite
diodiogod
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
17–29 / 29
Results: 29
No visual example yet
Explore the skilldiodiogod
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio…
No visual example yet
Explore the skillhuggingface
Build local voice agents with open-source models
No visual example yet
Explore the skilljianchang512
Voice Recognition to Text Tool / 一个离线运行的本地音视频转字幕工具,输出json、srt字幕、纯文字格式
No visual example yet
Explore the skillfikrikarim
On-device, real-time multimodal AI. Have natural voice and vision conversations with an AI that runs entirely on your machine. Powered by Gemma 4 E2B and Kokoro.
No visual example yet
Explore the skillhkchengrex
[CVPR 2025] MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
No visual example yet
Explore the skillmodal-labs
No visual example yet
Explore the skillk2-fsa
Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, Windows,…
No visual example yet
Explore the skillPluviobyte
Generate a controlled local narration workflow with auditions, version tracking, and subtitle-ready final audio.
No visual example yet
Explore the skillcalesthio
Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating audio content for videos, podcasts, or games. Triggers include generating voice…
No visual example yet
Explore the skillNVlabs
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
No visual example yet
Explore the skillGetStream
Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.
No visual example yet
Explore the skillzebbern
80+ free AI services for chat, image, video, voice & APIs (may sometimes include access to lead gen ai models for free)
No visual example yet
Explore the skillcalesthio
Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API. Use when: (1) Choosing a specific avatar and v…