video-voiceover
zenstory-ai
把带时间戳的 narration.json 合成为中文解说音频。使用 MiMo TTS(mimo-v2.5-tts)或 Fish Audio(s2.1-pro-free)逐段生成语音, 按时间窗动态适配语速并处理响度;输入输出时间线上的旁白,产出 tts_segments 与 tts_meta.json。 旧版直接剪辑路径也可显式传入…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
Page 27 · 16 shown · 744 public entries
Results: 744
zenstory-ai
把带时间戳的 narration.json 合成为中文解说音频。使用 MiMo TTS(mimo-v2.5-tts)或 Fish Audio(s2.1-pro-free)逐段生成语音, 按时间窗动态适配语速并处理响度;输入输出时间线上的旁白,产出 tts_segments 与 tts_meta.json。 旧版直接剪辑路径也可显式传入…
wildminder
ComfyUI custom node for the VibeVoice TTS. Expressive, long-form, multi-speaker conversational audio
githubharald
Connectionist Temporal Classification (CTC) decoder with dictionary and language model.
petermg
Modified version of Chatterbox that accepts text files as input and no character restrictions. I use it to make audiobooks, especially for my kids.
Azure-Samples
A simple example implementation of the VoiceRAG pattern to power interactive voice generative AI experiences using RAG with Azure AI Search and Azure OpenAI's gpt-4o-rea…
MarcosNahuel
Local NotebookLM for Claude Code via Google Antigravity (agy / Gemini 3.x): /agy:notebook turns a folder of documents into per-doc summaries + a relevance index + a cite…
theaiautomators
Open-source, self-hosted alternative to NotebookLM. Chat with your documents, generate audio summaries, and ground AI in your own sources—built with Supabase and N8N on…
cosmicstack-labs
Automated daily tech briefing — multi-source collection → knowledge-base deduplication → AI summarization → TTS speech synthesis, generating MP3 audio briefings
shibing624
Automatic Speech Recognition(ASR), Text-To-Speech(TTS) engine. 中英语音识别、多角色语音合成,支持多语言,准确率高
ID-LoRA
Custom ComfyUI node for generating videos with audio-visual identity based on a reference voice and image
ccoreilly
A speech recognition library running in the browser thanks to a WebAssembly build of Vosk
lucidrains
Implementation of E2-TTS, "Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS", in Pytorch
d3cod3
Mosaic, an openFrameworks/ImGui based Visual Patching Creative-Coding Platform
toverainc
Open source, local, and self-hosted highly optimized language inference server supporting ASR/STT, TTS, and LLM across WebRTC, REST, and WS
double22a
The dataset of Speech Recognition
egorsmkv
🇺🇦 Speech Recognition & Synthesis for Ukrainian