暂未收录效果图
查看技能说明Speech Recognition
Uberi
Speech recognition module for Python, supporting several engines and APIs, online and offline.
OPENAGENTSKILL / DIRECTORY
为下一项任务找到合适的技能。探索适用于 Codex、Claude Code、Cursor 等 Agent 的工具。
29 Skills
搜索结果: 29
暂未收录效果图
查看技能说明Uberi
Speech recognition module for Python, supporting several engines and APIs, online and offline.
暂未收录效果图
查看技能说明m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
暂未收录效果图
查看技能说明TalAter
暂未收录效果图
查看技能说明cmusphinx
暂未收录效果图
查看技能说明MahmoudAshraf97
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
暂未收录效果图
查看技能说明flashlight
暂未收录效果图
查看技能说明jianchang512
Voice Recognition to Text Tool / 一个离线运行的本地音视频转字幕工具,输出json、srt字幕、纯文字格式
暂未收录效果图
查看技能说明huggingface
Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.
暂未收录效果图
查看技能说明alphacep
暂未收录效果图
查看技能说明ageitgey
The world's simplest facial recognition api for Python and the command line
暂未收录效果图
查看技能说明babysor
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
暂未收录效果图
查看技能说明rany2
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
暂未收录效果图
查看技能说明espeak-ng
eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.
暂未收录效果图
查看技能说明KoljaB
暂未收录效果图
查看技能说明OpenMOSS
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and…
暂未收录效果图
查看技能说明jaywalnut310
VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech