ChatTTS
2noise
A generative speech model for daily dialogue.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
97–112 / 115
Results: 115
2noise
A generative speech model for daily dialogue.
xorbitsai
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all throug…
modelscope
Open-source, accurate and easy-to-use video speech recognition & clipping tool. LLM-based AI clipping integrated.
instillai
:speech_balloon: Machine Learning Course with Python:
linto-ai
Multilingual Automatic Speech Recognition with word-level timestamps and confidence
kubeai-project
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
nexu-io
Audio generation skill — jingles, beds, voiceover, and sound effects. Routes music requests to Suno V5 / Udio / Lyria, speech to MiniMax TTS / FishAudio / ElevenLabs V3,…
nazdridoy
A CLI text-to-speech tool using the Kokoro model, supporting multiple languages, voices (with blending), and various input formats including EPUB books and PDF documents.
calesthio
DashScope (Alibaba Cloud Bailian / 阿里云百炼) integration — image generation (qwen-image-2.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen…
calesthio
Generate Mandarin and multilingual narration with Volcengine Doubao Speech 2.0. Use when creating Chinese voiceovers, when the user prefers Doubao/Volcengine/火山引擎/豆包 TTS…
byjlw
Analyze videos using LLMs, Computer Vision and Automatic Speech Recognition
openclaw
Local speech-to-text with the Whisper CLI (no API key).
alexpinel
Text-To-Speech, RAG, and LLMs. All local!
LucasBassetti
:speech_balloon: Easy way to create conversation chats
TimoBolkart
This codebase demonstrates how to synthesize realistic 3D character animations given an arbitrary speech signal and a static character mesh.
MengTo
Generate ElevenLabs text-to-speech audio from scripts or inline text using local voice profiles. Use when the user asks for ElevenLabs, text-to-speech, TTS, narration, v…