Mediapipe
google-ai-edge
Cross-platform, customizable ML solutions for live and streaming media.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–16 / 101
Results: 101
google-ai-edge
Cross-platform, customizable ML solutions for live and streaming media.
NanmiCoder
小红书笔记 | 评论爬虫、抖音视频 | 评论爬虫、快手视频 | 评论爬虫、B站视频 | 评论爬虫、微博帖子 | 评论爬虫、百度贴吧帖子 | 知乎问答文章 | 评论爬虫。支持多平台社交媒体内容抓取,提供完整的数据采集解决方案。
RVC-Boss
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
biolab
🍊 :bar_chart: :bulb: Orange: Interactive data analysis
obss
Framework agnostic sliced/tiled inference + interactive ui + error analysis plots
babysor
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
ml-tooling
🪄 Turns your machine learning code into microservices with web API, interactive GUI, and more.
unslothai
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
ggml-org
Port of OpenAI's Whisper model in C/C++
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
huggingface
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
OpenBMB
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
invoke-ai
Invoke is a leading creative engine for Stable Diffusion models, empowering professionals, artists, and enthusiasts to generate and create visual media using the latest…
AIDC-AI
🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine
Tencent-Hunyuan
HunyuanVideo: A Systematic Framework For Large Video Generation Model