Vosk API
alphacep
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 319
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 319
alphacep
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
PaddlePaddle
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System,…
leandromoreira
FFmpeg libav tutorial - learn how media works from basic to transmuxing, transcoding and more. Translations: 🇺🇸 🇨🇳 🇰🇷 🇪🇸 🇻🇳 🇧🇷 🇷🇺
nyanmisaka
FFmpeg with async and zero-copy Rockchip MPP & RGA support
ErickWendel
JS Expert Week 8.0 - 🎥Pre processing videos before uploading in the browser 😏
jik876
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
kan-bayashi
Unofficial Parallel WaveGAN (+ MelGAN & Multi-band MelGAN & HiFi-GAN & StyleMelGAN) with Pytorch
jishengpeng
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
tmchow
illo skill — an AI agent skill that turns ideas and articles into original print-style editorial illustrations, starring a recurring mascot. 30+ characters packs, with a…
leeguooooo
Use your ChatGPT subscription to generate images from the command line — no OPENAI_API_KEY, no gateway, no daemon. Zero-dep Python CLI + AI-agent skill.
wanshuiyin
Bilingual (中文+EN) ML / LLM / diffusion / agent interview cheat sheets for AI 秋招 — generated by ARIS /interview-cheatsheet, rendered by /render-html into single-file HTML…
RVC-Boss
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
babysor
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
FoundationVision
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation…