Mediapipe
google-ai-edge
Cross-platform, customizable ML solutions for live and streaming media.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
65–80 / 481
Results: 481
google-ai-edge
Cross-platform, customizable ML solutions for live and streaming media.
emilkowalski
Reverse-lookup glossary that turns a vague description of a web animation or motion effect into its exact term ("the bouncy thing when a popover opens" → Pop in; "the iO…
Developer-Y
List of Computer Science courses with video lectures.
unslothai
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
ggml-org
Port of OpenAI's Whisper model in C/C++
AaronFeng753
Video, Image and GIF upscale/enlarge(Super-Resolution) and Video frame interpolation. Achieved with Waifu2x, Real-ESRGAN, Real-CUGAN, RTX Video Super Resolution VSR, SRM…
alphacep
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
waooAI
首家工业级全流程 AI 影视生产平台。Industry-first professional AI Agent platform for controllable film & video production. From shorts to live-action with Hollywood-standard workflows.
huggingface
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
OpenBMB
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
invoke-ai
Invoke is a leading creative engine for Stable Diffusion models, empowering professionals, artists, and enthusiasts to generate and create visual media using the latest…
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
AIDC-AI
🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine
browser-use
Edit videos with coding agents
FunAudioLLM
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
Design direction, Figma implementation, UI review, React performance, browser QA, and safe preview deployments.
Code review, repo inspection, testing, planning, shipping, and engineering workflows for coding agents.
Research, recent-events briefings, PDF parsing, markdown conversion, RAG ingestion, and knowledge workflows.
Stock analysis, market research, quant backtesting, financial data, and investment research skills.
Crawling, scraping, extraction, browser automation, structured data capture, and website-to-markdown workflows.
Presentation generation, editable PPTX, slide decks, speaker notes, and visual storytelling workflows.
Video-generation prompts, B-roll, Vox-style explainers, camera direction, captions, and creative-production workflows.
Image, video, creative production, UI design, multimodal generation, and visual workflow skills.