WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
$ npx skills add m-bain/whisperXScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
AI Agent Skill Repository
Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.
Live registry search
Exact name and slug matches are checked against the live registry before ranked alternatives.
Decision filters
Showing 1-13 of 13 ranked candidates matching "speech-text-pretraining"
Best blend of relevance, quality, freshness, and verified outcomes
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
$ npx skills add m-bain/whisperXScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Port of OpenAI's Whisper model in C/C++
$ npx skills add ggml-org/whisper.cppScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
OpenAI Agents + CLI · 4 targets
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
$ npx skills add alphacep/vosk-apiScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
$ npx skills add unslothai/unslothScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
OpenAI Agents + CLI · 4 targets
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
$ npx skills add RVC-Boss/GPT-SoVITSScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
$ npx skills add OpenBMB/VoxCPMScenario Design and creative · I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
CLI + Codex · 4 targets
Rich is a Python library for rich text and beautiful formatting in the terminal.
$ npx skills add Textualize/richScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
A lightning-fast search engine API bringing AI-powered hybrid search to your sites and applications.
$ npx skills add meilisearch/meilisearchScenario RAG and knowledge · I need my agent to build a RAG workflow over documents and retrieve reliable context.
CLI + Codex · 4 targets
Lightning-fast and Powerful Code Editor written in Rust
$ npx skills add lapce/lapceScenario Coding agents · I need a coding agent that can understand a repository, edit code, and review pull requests.
CLI + Codex · 4 targets
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and traini…
$ npx skills add huggingface/transformersScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
ShareX is a free and open-source application that enables users to capture or record any area of their screen with a single keystroke. It also supports uploading images,…
$ npx skills add ShareX/ShareXScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
Write precise Seedance 2.0 prompts for multimodal video, camera movement, editing, music, and product storytelling.
$ npx skills add dexhunter/seedance2-skill --skill seedance-prompt-enScenario Video creation · I need my agent to turn a topic or product into a short video, generate B-roll, write video prompts, and prepare a…
Claude Code + OpenAI Agents · 4 targets
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
$ npx skills add ocrmypdf/OCRmyPDFScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
AI Agent Skill Repository
Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.
Live registry search
Exact name and slug matches are checked against the live registry before ranked alternatives.
Decision filters
Showing 1-13 of 13 ranked candidates matching "speech-text-pretraining"
Best blend of relevance, quality, freshness, and verified outcomes
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
$ npx skills add m-bain/whisperXScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Port of OpenAI's Whisper model in C/C++
$ npx skills add ggml-org/whisper.cppScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
OpenAI Agents + CLI · 4 targets
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
$ npx skills add alphacep/vosk-apiScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
$ npx skills add unslothai/unslothScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
OpenAI Agents + CLI · 4 targets
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
$ npx skills add RVC-Boss/GPT-SoVITSScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
$ npx skills add OpenBMB/VoxCPMScenario Design and creative · I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
CLI + Codex · 4 targets
Rich is a Python library for rich text and beautiful formatting in the terminal.
$ npx skills add Textualize/richScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
A lightning-fast search engine API bringing AI-powered hybrid search to your sites and applications.
$ npx skills add meilisearch/meilisearchScenario RAG and knowledge · I need my agent to build a RAG workflow over documents and retrieve reliable context.
CLI + Codex · 4 targets
Lightning-fast and Powerful Code Editor written in Rust
$ npx skills add lapce/lapceScenario Coding agents · I need a coding agent that can understand a repository, edit code, and review pull requests.
CLI + Codex · 4 targets
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and traini…
$ npx skills add huggingface/transformersScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
ShareX is a free and open-source application that enables users to capture or record any area of their screen with a single keystroke. It also supports uploading images,…
$ npx skills add ShareX/ShareXScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets
Write precise Seedance 2.0 prompts for multimodal video, camera movement, editing, music, and product storytelling.
$ npx skills add dexhunter/seedance2-skill --skill seedance-prompt-enScenario Video creation · I need my agent to turn a topic or product into a short video, generate B-roll, write video prompts, and prepare a…
Claude Code + OpenAI Agents · 4 targets
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
$ npx skills add ocrmypdf/OCRmyPDFScenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
CLI + Codex · 4 targets