Wenet
wenet-e2e
Production First and Production Ready End-to-End Speech Recognition Toolkit
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 319
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 319
wenet-e2e
Production First and Production Ready End-to-End Speech Recognition Toolkit
kane50613
Render JSX, HTML, and CSS to images without a headless browser. OG cards, animated GIFs, and video frames from Node.js, edge runtimes, browsers, or Rust. Drop-in next/og…
s60sc
ESP32 Camera motion capture application to record JPEGs to SD card as AVI files and stream to browser as MJPEG. If a microphone is installed then a WAV file is also crea…
ErickWendel
JS Expert Week 8.0 - 🎥Pre processing videos before uploading in the browser 😏
Uminosachi
Inpaint Anything extension performs stable diffusion inpainting on a browser UI using masks from Segment Anything.
tejaswigowda
A browser-based video editor powered by ffmpeg.wasm. No uploads, no servers -- all processing happens locally in your browser using WebAssembly.
cboard-org
Augmentative and Alternative Communication (AAC) system with text-to-speech for the browser
pnlpal
📚 A customizable dictionary extension that supports double-click lookups in 20+ languages, 1000+ dictionaries, text-to-speech, translation and Anki integration.
toki-plus
全自动短视频搬运工具,支持自动下载、去重、AI生成标题+标签、上传,可二开扩展至多平台,例如:TikTok->视频号/抖音/小红书、抖音->TikTok/视频号/小红书......video-processing, automation, tiktok, selenium, pyqt5, ffmpeg, bot, data-scrapi…
deepgram
Official Python SDK for Deepgram.
invoke-ai
Invoke is a leading creative engine for Stable Diffusion models, empowering professionals, artists, and enthusiasts to generate and create visual media using the latest…
AIDC-AI
🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine
coqui-ai
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
RVC-Boss
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)