Whisper Diarization
MahmoudAshraf97
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Search retrieves candidates across the registry and ranks a bounded shortlist by task fit. This count is matching candidates, not the registry total. No suitable match? Try a specific tool or task.
1–13 / 13
Results: 13
MahmoudAshraf97
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
transcriptionstream
turnkey self-hosted offline transcription and diarization service with llm summary
izwi-ai
Voice AI runtime. Local first transcription, speaker diarization, TTS, and voice cloning with an OpenAI compatible API.
k2-fsa
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Su…
m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
FunAudioLLM
End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.
PaddlePaddle
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System,…
OpenMOSS
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and…
Enemyx-net
A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly within your C…
soniqo
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
Saganaki22
OmniVoice TTS nodes for ComfyUI - Zero-shot multilingual text-to-speech with voice cloning, voice design, and multi-speaker dialogue
wildminder
ComfyUI custom node for the VibeVoice TTS. Expressive, long-form, multi-speaker conversational audio
SamirPaulb
A desktop application that uses AI to translate voice between languages in real time, while preserving the speaker's tone and emotion.
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.