Dia2
nari-labs
TTS model capable of streaming conversational audio in realtime.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
Page 12 · 16 shown · 743 public entries
Results: 743
nari-labs
TTS model capable of streaming conversational audio in realtime.
metavoiceio
Foundational model for human-like, expressive TTS
sooftware
[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)
huggingface
Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.
TensorSpeech
:stuck_out_tongue_closed_eyes: TensorFlowTTS: Real-Time State-of-the-art Speech Synthesis for Tensorflow 2 (supported including English, French, Korean, Chinese, German…
agan-j
小牛视频翻译 是一款支持本地视频翻译、字幕翻译和 YouTube 视频翻译下载的 AI 工具,集成自动语音识别与多语言翻译功能,助力创作者高效完成视频翻译,应用于视频本地化与视频出海场景。
alphacep
Offline speech recognition for Android with Vosk library.
RefoundAI
86 product management skills from Lenny's Podcast for Claude Code and AI agents. Hiring, user research, strategy, shipping, and more.
Saik0s
The open-source iOS app that's making quality voice transcription more accessible on mobile devices.
Agions
scene-fab - AI 影视解说创作工具 | 智能拆条 · AI 解说生成 · 一键配音合成
pykaldi
A Python wrapper for Kaldi
k2-fsa
Speech-to-text server framework with next-gen Kaldi
BinWang28
The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
Vonage
Vonage REST API client for PHP. API support for SMS, Voice, Text-to-Speech, Numbers, Verify (2FA) and more.
sandrohanea
Whisper.net. Speech to text made simple using Whisper Models
calesthio
Generate expressive, multilingual narration with fish.audio (S1 / S2-generation models) and reuse cloned voices via reference_id. Use when the user prefers fish.audio/Fi…