Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
$ npx skills add microsoft/SpeechT5Scenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
AI Agent Skill Repository
Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.
Live registry search
Exact name and slug matches are checked against the live registry before ranked alternatives.
Decision filters
Showing 1-16 of 121 ranked candidates matching "speech-pretraining"
Best blend of relevance, quality, freshness, and verified outcomes
Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
$ npx skills add microsoft/SpeechT5Scenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System,…
$ npx skills add PaddlePaddle/PaddleSpeechScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
A PyTorch-based Speech Toolkit
$ npx skills add speechbrain/speechbrainScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Speech recognition module for Python, supporting several engines and APIs, online and offline.
$ npx skills add Uberi/speech_recognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
A Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统
$ npx skills add nl8590687/ASRT_SpeechRecognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow
$ npx skills add zzw922cn/Automatic_Speech_RecognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks
$ npx skills add pannous/tensorflow-speech-recognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
The neural network model is capable of detecting five different male/female emotions from audio speeches. (Deep Learning, NLP, Python)
$ npx skills add MiteshPuthran/Speech-Emotion-AnalyzerScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
$ npx skills add coqui-ai/open-speech-corporaScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
$ npx skills add ictnlp/StreamSpeechScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
$ npx skills add soniqo/speech-swiftScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Speech Recognition for React Native Expo projects
$ npx skills add jamsch/expo-speech-recognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Build local voice agents with open-source models
$ npx skills add huggingface/speech-to-speechScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.
$ npx skills add lenML/Speech-AI-ForgeScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Microsoft Text-to-Speech API sample code in several languages, part of Cognitive Services.
$ npx skills add Azure-Samples/Cognitive-Speech-TTSScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Tools for handling multimodal data in machine learning projects.
$ npx skills add lhotse-speech/lhotseScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Page 1
Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.
Try the agent resolve APIAI Agent Skill Repository
Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.
Live registry search
Exact name and slug matches are checked against the live registry before ranked alternatives.
Decision filters
Showing 1-16 of 121 ranked candidates matching "speech-pretraining"
Best blend of relevance, quality, freshness, and verified outcomes
Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
$ npx skills add microsoft/SpeechT5Scenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System,…
$ npx skills add PaddlePaddle/PaddleSpeechScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
A PyTorch-based Speech Toolkit
$ npx skills add speechbrain/speechbrainScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Speech recognition module for Python, supporting several engines and APIs, online and offline.
$ npx skills add Uberi/speech_recognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
A Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统
$ npx skills add nl8590687/ASRT_SpeechRecognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow
$ npx skills add zzw922cn/Automatic_Speech_RecognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks
$ npx skills add pannous/tensorflow-speech-recognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
The neural network model is capable of detecting five different male/female emotions from audio speeches. (Deep Learning, NLP, Python)
$ npx skills add MiteshPuthran/Speech-Emotion-AnalyzerScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
$ npx skills add coqui-ai/open-speech-corporaScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
$ npx skills add ictnlp/StreamSpeechScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
$ npx skills add soniqo/speech-swiftScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Speech Recognition for React Native Expo projects
$ npx skills add jamsch/expo-speech-recognitionScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Build local voice agents with open-source models
$ npx skills add huggingface/speech-to-speechScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.
$ npx skills add lenML/Speech-AI-ForgeScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Microsoft Text-to-Speech API sample code in several languages, part of Cognitive Services.
$ npx skills add Azure-Samples/Cognitive-Speech-TTSScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Tools for handling multimodal data in machine learning projects.
$ npx skills add lhotse-speech/lhotseScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Page 1
Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.
Try the agent resolve API