暂未收录效果图
查看技能说明Speech Recognition
Uberi
Speech recognition module for Python, supporting several engines and APIs, online and offline.
OPENAGENTSKILL / DIRECTORY
为下一项任务找到合适的技能。探索适用于 Codex、Claude Code、Cursor 等 Agent 的工具。
21 Skills
搜索结果: 21
暂未收录效果图
查看技能说明Uberi
Speech recognition module for Python, supporting several engines and APIs, online and offline.
暂未收录效果图
查看技能说明m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
暂未收录效果图
查看技能说明ageitgey
The world's simplest facial recognition api for Python and the command line
暂未收录效果图
查看技能说明MahmoudAshraf97
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
暂未收录效果图
查看技能说明flashlight
暂未收录效果图
查看技能说明TalAter
暂未收录效果图
查看技能说明cmusphinx
暂未收录效果图
查看技能说明jianchang512
Voice Recognition to Text Tool / 一个离线运行的本地音视频转字幕工具,输出json、srt字幕、纯文字格式
暂未收录效果图
查看技能说明babysor
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
暂未收录效果图
查看技能说明ThioJoe
Automatically translates the text of a video based on a subtitle file, and then uses AI voice services to create a new dubbed & translated audio track where the speech i…
暂未收录效果图
查看技能说明rany2
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
暂未收录效果图
查看技能说明espeak-ng
eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.
暂未收录效果图
查看技能说明KoljaB
暂未收录效果图
查看技能说明OpenMOSS
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and…
暂未收录效果图
查看技能说明nateshmbhat
暂未收录效果图
查看技能说明Aratako
A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control