Cross-platform, customizable ML solutions for live and streaming media.
$ npx skills add google-ai-edge/mediapipeScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
AI Agent Skill Repository
Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.
Live registry search
Exact name and slug matches are checked against the live registry before ranked alternatives.
Decision filters
Showing 1-16 of 149 ranked candidates matching "media-management"
Best blend of relevance, quality, freshness, and verified outcomes
Cross-platform, customizable ML solutions for live and streaming media.
$ npx skills add google-ai-edge/mediapipeScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
小红书笔记 | 评论爬虫、抖音视频 | 评论爬虫、快手视频 | 评论爬虫、B站视频 | 评论爬虫、微博帖子 | 评论爬虫、百度贴吧帖子 | 知乎问答文章 | 评论爬虫。支持多平台社交媒体内容抓取,提供完整的数据采集解决方案。
$ npx skills add NanmiCoder/MediaCrawlerScenario Coding agents · I need a coding agent that can understand a repository, edit code, and review pull requests.
CLI + Codex · 4 targets
Bring projects, wikis, and teams together with AI. AppFlowy is the AI collaborative workspace where you achieve more without losing control of your data. The leading ope…
$ npx skills add AppFlowy-IO/AppFlowyScenario RAG and knowledge · I need my agent to build a RAG workflow over documents and retrieve reliable context.
CLI + Codex · 4 targets
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
$ npx skills add m-bain/whisperXScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
$ npx skills add unslothai/unslothScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
OpenAI Agents + CLI · 4 targets
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
$ npx skills add RVC-Boss/GPT-SoVITSScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Port of OpenAI's Whisper model in C/C++
$ npx skills add ggml-org/whisper.cppScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
OpenAI Agents + CLI · 4 targets
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
$ npx skills add FunAudioLLM/SenseVoiceScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
$ npx skills add huggingface/diffusersScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
$ npx skills add OpenBMB/VoxCPMScenario Design and creative · I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
CLI + Codex · 4 targets
Invoke is a leading creative engine for Stable Diffusion models, empowering professionals, artists, and enthusiasts to generate and create visual media using the latest…
$ npx skills add invoke-ai/InvokeAIScenario Design and creative · I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
CLI + Codex · 4 targets
Stable Diffusion web UI
$ npx skills add AUTOMATIC1111/stable-diffusion-webuiScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine
$ npx skills add AIDC-AI/Pixelle-VideoScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Video, Image and GIF upscale/enlarge(Super-Resolution) and Video frame interpolation. Achieved with Waifu2x, Real-ESRGAN, Real-CUGAN, RTX Video Super Resolution VSR, SRM…
$ npx skills add AaronFeng753/Waifu2x-Extension-GUIScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
$ npx skills add index-tts/index-ttsScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
$ npx skills add alphacep/vosk-apiScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Page 1
Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.
Try the agent resolve APIAI Agent Skill Repository
Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.
Live registry search
Exact name and slug matches are checked against the live registry before ranked alternatives.
Decision filters
Showing 1-16 of 149 ranked candidates matching "media-management"
Best blend of relevance, quality, freshness, and verified outcomes
Cross-platform, customizable ML solutions for live and streaming media.
$ npx skills add google-ai-edge/mediapipeScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
小红书笔记 | 评论爬虫、抖音视频 | 评论爬虫、快手视频 | 评论爬虫、B站视频 | 评论爬虫、微博帖子 | 评论爬虫、百度贴吧帖子 | 知乎问答文章 | 评论爬虫。支持多平台社交媒体内容抓取,提供完整的数据采集解决方案。
$ npx skills add NanmiCoder/MediaCrawlerScenario Coding agents · I need a coding agent that can understand a repository, edit code, and review pull requests.
CLI + Codex · 4 targets
Bring projects, wikis, and teams together with AI. AppFlowy is the AI collaborative workspace where you achieve more without losing control of your data. The leading ope…
$ npx skills add AppFlowy-IO/AppFlowyScenario RAG and knowledge · I need my agent to build a RAG workflow over documents and retrieve reliable context.
CLI + Codex · 4 targets
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
$ npx skills add m-bain/whisperXScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
$ npx skills add unslothai/unslothScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
OpenAI Agents + CLI · 4 targets
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
$ npx skills add RVC-Boss/GPT-SoVITSScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Port of OpenAI's Whisper model in C/C++
$ npx skills add ggml-org/whisper.cppScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
OpenAI Agents + CLI · 4 targets
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
$ npx skills add FunAudioLLM/SenseVoiceScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
$ npx skills add huggingface/diffusersScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
$ npx skills add OpenBMB/VoxCPMScenario Design and creative · I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
CLI + Codex · 4 targets
Invoke is a leading creative engine for Stable Diffusion models, empowering professionals, artists, and enthusiasts to generate and create visual media using the latest…
$ npx skills add invoke-ai/InvokeAIScenario Design and creative · I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
CLI + Codex · 4 targets
Stable Diffusion web UI
$ npx skills add AUTOMATIC1111/stable-diffusion-webuiScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine
$ npx skills add AIDC-AI/Pixelle-VideoScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Video, Image and GIF upscale/enlarge(Super-Resolution) and Video frame interpolation. Achieved with Waifu2x, Real-ESRGAN, Real-CUGAN, RTX Video Super Resolution VSR, SRM…
$ npx skills add AaronFeng753/Waifu2x-Extension-GUIScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
$ npx skills add index-tts/index-ttsScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
$ npx skills add alphacep/vosk-apiScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Page 1
Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.
Try the agent resolve API