Cross-platform, customizable ML solutions for live and streaming media.
$ npx skills add google-ai-edge/mediapipeScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
AI Agent Skill Repository
Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.
Live registry search
Exact name and slug matches are checked against the live registry before ranked alternatives.
Decision filters
Showing 1-16 of 44 ranked candidates matching "media-library"
Best blend of relevance, quality, freshness, and verified outcomes
Cross-platform, customizable ML solutions for live and streaming media.
$ npx skills add google-ai-edge/mediapipeScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
$ npx skills add m-bain/whisperXScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
$ npx skills add unslothai/unslothScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
OpenAI Agents + CLI · 4 targets
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
$ npx skills add RVC-Boss/GPT-SoVITSScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Port of OpenAI's Whisper model in C/C++
$ npx skills add ggml-org/whisper.cppScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
OpenAI Agents + CLI · 4 targets
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
$ npx skills add huggingface/diffusersScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
$ npx skills add OpenBMB/VoxCPMScenario Design and creative · I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
CLI + Codex · 4 targets
Video, Image and GIF upscale/enlarge(Super-Resolution) and Video frame interpolation. Achieved with Waifu2x, Real-ESRGAN, Real-CUGAN, RTX Video Super Resolution VSR, SRM…
$ npx skills add AaronFeng753/Waifu2x-Extension-GUIScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
$ npx skills add alphacep/vosk-apiScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
$ npx skills add deepspeedai/DeepSpeedScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Apache ECharts is a powerful, interactive charting and data visualization library for browser
$ npx skills add apache/echartsScenario Data analysis · I need my agent to analyze CSV data, produce insights, and explain trends.
Browser agents + CLI · 4 targets
Agent-led recent-trends research across social, prediction markets, video, code, and the web.
$ npx skills add mvanhorn/last30days-skill -gScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
Claude Code + OpenAI Agents · 4 targets
An Open Source Machine Learning Framework for Everyone
$ npx skills add tensorflow/tensorflowScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and traini…
$ npx skills add huggingface/transformersScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Tensors and Dynamic neural networks in Python with strong GPU acceleration
$ npx skills add pytorch/pytorchScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
$ npx skills add rasbt/LLMs-from-scratchScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
OpenAI Agents + CLI · 4 targets
Page 1
Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.
Try the agent resolve APIAI Agent Skill Repository
Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.
Live registry search
Exact name and slug matches are checked against the live registry before ranked alternatives.
Decision filters
Showing 1-16 of 44 ranked candidates matching "media-library"
Best blend of relevance, quality, freshness, and verified outcomes
Cross-platform, customizable ML solutions for live and streaming media.
$ npx skills add google-ai-edge/mediapipeScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
$ npx skills add m-bain/whisperXScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
$ npx skills add unslothai/unslothScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
OpenAI Agents + CLI · 4 targets
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
$ npx skills add RVC-Boss/GPT-SoVITSScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Port of OpenAI's Whisper model in C/C++
$ npx skills add ggml-org/whisper.cppScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
OpenAI Agents + CLI · 4 targets
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
$ npx skills add huggingface/diffusersScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
$ npx skills add OpenBMB/VoxCPMScenario Design and creative · I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
CLI + Codex · 4 targets
Video, Image and GIF upscale/enlarge(Super-Resolution) and Video frame interpolation. Achieved with Waifu2x, Real-ESRGAN, Real-CUGAN, RTX Video Super Resolution VSR, SRM…
$ npx skills add AaronFeng753/Waifu2x-Extension-GUIScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
$ npx skills add alphacep/vosk-apiScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
$ npx skills add deepspeedai/DeepSpeedScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Apache ECharts is a powerful, interactive charting and data visualization library for browser
$ npx skills add apache/echartsScenario Data analysis · I need my agent to analyze CSV data, produce insights, and explain trends.
Browser agents + CLI · 4 targets
Agent-led recent-trends research across social, prediction markets, video, code, and the web.
$ npx skills add mvanhorn/last30days-skill -gScenario Research agents · I need my agent to research a topic, compare sources, and produce a concise report.
Claude Code + OpenAI Agents · 4 targets
An Open Source Machine Learning Framework for Everyone
$ npx skills add tensorflow/tensorflowScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and traini…
$ npx skills add huggingface/transformersScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Tensors and Dynamic neural networks in Python with strong GPU acceleration
$ npx skills add pytorch/pytorchScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
$ npx skills add rasbt/LLMs-from-scratchScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
OpenAI Agents + CLI · 4 targets
Page 1
Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.
Try the agent resolve API