StreamSpeech
ictnlp
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
65–80 / 96
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 96
ictnlp
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
sdkcarlos
A voice control - voice commands - speech recognition and speech synthesis javascript library. Create your own siri,google now or cortana with Google Chrome within your…
alphacep
WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries
mravanelli
SincNet is a neural architecture for efficiently processing raw audio samples.
alumae
Real-time full-duplex speech recognition server, based on the Kaldi toolkit and the GStreamer framwork.
TensorSpeech
:zap: TensorFlowASR: Almost State-of-the-art Automatic Speech Recognition in Tensorflow 2. Supported languages that can use characters or subwords
alesaccoia
Near-Realtime audio transcription using self-hosted Whisper and WebSocket in Python/JS
k2-fsa
Speech-to-text server framework with next-gen Kaldi
BinWang28
The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
sandrohanea
Whisper.net. Speech to text made simple using Whisper Models
soniqo
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
TheStageAI
Optimized Whisper models for streaming and on-device use
AutoArk
[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!
peteonrails
Voice-to-text with push-to-talk for Wayland compositors