WhisperX
m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
17–32 / 108
Results: 108
m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
babysor
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
ggml-org
Port of OpenAI's Whisper model in C/C++
AIDC-AI
🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine
Tencent-Hunyuan
HunyuanVideo: A Systematic Framework For Large Video Generation Model
KlingAIResearch
jianchang512
Translate the video from one language to another and embed dubbing & subtitles.
Wan-Video
Wan: Open and Advanced Large-Scale Video Generative Models
openvinotoolkit
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
Zulko
HKUDS
"ViMax: Agentic Video Generation (Director, Screenwriter, Producer, and Video Generator All-in-One)"
nari-labs
A TTS model capable of generating ultra-realistic dialogue in one pass.
Uberi
Speech recognition module for Python, supporting several engines and APIs, online and offline.
rany2
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
VectorSpaceLab
OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
TalAter
💬 Speech recognition for your site