Irodori TTS
Aratako
A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
17–32 / 84
Results: 84
Aratako
A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control
lhotse-speech
Tools for handling multimodal data in machine learning projects.
open-mmlab
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get…
TensorSpeech
:stuck_out_tongue_closed_eyes: TensorFlowTTS: Real-Time State-of-the-art Speech Synthesis for Tensorflow 2 (supported including English, French, Korean, Chinese, German…
m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
MoonInTheRiver
DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (SVS & TTS); AAAI 2022; Official code
calesthio
Generate neural narration audio using Azure AI Speech (REST text-to-speech). Use when synthesizing voiceovers or narration in OpenMontage. Optional cloud TTS provider —…
calesthio
Transcribe audio to text using Azure AI Speech (Fast Transcription REST API). Use when converting audio/video to text, generating subtitles, or processing spoken content…
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
alphacep
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
wenet-e2e
Production First and Production Ready End-to-End Speech Recognition Toolkit
MahmoudAshraf97
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
flashlight
Facebook AI Research's Automatic Speech Recognition Toolkit
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.