Build local voice agents with open-source models
$ npx skills add huggingface/speech-to-speechScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
AI Agent Skill Repository
Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.
Live registry search
Exact name and slug matches are checked against the live registry before ranked alternatives.
Decision filters
Showing 1-16 of 119 ranked candidates matching "conversational-speech-synthesis"
Best blend of relevance, quality, freshness, and verified outcomes
Build local voice agents with open-source models
$ npx skills add huggingface/speech-to-speechScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System,…
$ npx skills add PaddlePaddle/PaddleSpeechScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
A PyTorch-based Speech Toolkit
$ npx skills add speechbrain/speechbrainScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Speech recognition module for Python, supporting several engines and APIs, online and offline.
$ npx skills add Uberi/speech_recognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
A Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统
$ npx skills add nl8590687/ASRT_SpeechRecognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow
$ npx skills add zzw922cn/Automatic_Speech_RecognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks
$ npx skills add pannous/tensorflow-speech-recognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
$ npx skills add soniqo/speech-swiftScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Speech Recognition for React Native Expo projects
$ npx skills add jamsch/expo-speech-recognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
End-to-End Speech Processing Toolkit
$ npx skills add espnet/espnetScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
$ npx skills add Blaizzy/mlx-audioScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.
$ npx skills add lenML/Speech-AI-ForgeScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Microsoft Text-to-Speech API sample code in several languages, part of Cognitive Services.
$ npx skills add Azure-Samples/Cognitive-Speech-TTSScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Tools for handling multimodal data in machine learning projects.
$ npx skills add lhotse-speech/lhotseScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
$ npx skills add rany2/edge-ttsScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get…
$ npx skills add open-mmlab/AmphionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Page 1
Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.
Try the agent resolve APIAI Agent Skill Repository
Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.
Live registry search
Exact name and slug matches are checked against the live registry before ranked alternatives.
Decision filters
Showing 1-16 of 119 ranked candidates matching "conversational-speech-synthesis"
Best blend of relevance, quality, freshness, and verified outcomes
Build local voice agents with open-source models
$ npx skills add huggingface/speech-to-speechScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System,…
$ npx skills add PaddlePaddle/PaddleSpeechScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
A PyTorch-based Speech Toolkit
$ npx skills add speechbrain/speechbrainScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Speech recognition module for Python, supporting several engines and APIs, online and offline.
$ npx skills add Uberi/speech_recognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
A Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统
$ npx skills add nl8590687/ASRT_SpeechRecognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow
$ npx skills add zzw922cn/Automatic_Speech_RecognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks
$ npx skills add pannous/tensorflow-speech-recognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
$ npx skills add soniqo/speech-swiftScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Speech Recognition for React Native Expo projects
$ npx skills add jamsch/expo-speech-recognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
End-to-End Speech Processing Toolkit
$ npx skills add espnet/espnetScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
$ npx skills add Blaizzy/mlx-audioScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.
$ npx skills add lenML/Speech-AI-ForgeScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Microsoft Text-to-Speech API sample code in several languages, part of Cognitive Services.
$ npx skills add Azure-Samples/Cognitive-Speech-TTSScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Tools for handling multimodal data in machine learning projects.
$ npx skills add lhotse-speech/lhotseScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
$ npx skills add rany2/edge-ttsScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get…
$ npx skills add open-mmlab/AmphionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Page 1
Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.
Try the agent resolve API