Sherpa
k2-fsa
Speech-to-text server framework with next-gen Kaldi
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
353–368 / 478
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 478
k2-fsa
Speech-to-text server framework with next-gen Kaldi
BinWang28
The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
sandrohanea
Whisper.net. Speech to text made simple using Whisper Models
soniqo
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
TheStageAI
Optimized Whisper models for streaming and on-device use
EDDiscovery
Captains log and 3d star map for Elite Dangerous
AutoArk
[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!
peteonrails
Voice-to-text with push-to-talk for Wayland compositors
thu-ml
Official implementation for "RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers" (ICML 2025) , UltraViCo (ICLR 2026) and UltraImage
tejaswigowda
A browser-based video editor powered by ffmpeg.wasm. No uploads, no servers -- all processing happens locally in your browser using WebAssembly.
VRCWizard
Speech to Text to Speech. Song now playing. Sends text as OSC messages to VRChat to display on avatar. (STTTS) (Speech to TTS) (VRC STT System) (VTuber TTS)
kandinskylab
Kandinsky 5.0: A family of diffusion models for Video & Image generation
cboard-org
Augmentative and Alternative Communication (AAC) system with text-to-speech for the browser
Picovoice
On-device Speech-to-Intent engine powered by deep learning
fluxions-ai
Real-time voice assistant — WebRTC streaming, faster-whisper ASR, local LLM, Vui Nano (300M) TTS. OpenAI Realtime API compatible. Voice cloning, barge-in, ~9× realtime o…