WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
每个推荐都保留与其仓库、审计和安装路径的明确关联。
搜索结果: voice-recognition
英文目录1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
The world's simplest facial recognition api for Python and the command line
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
A Lightweight Face Recognition and Facial Attribute Analysis (Age, Gender, Emotion and Race) Library for Python
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
💬 Open source machine learning framework to automate text- and voice-based conversations: NLU, dialogue management, connect to Slack, Facebook, and more - Create chatbots and voice assistants
🌈一个跨平台的划词翻译和OCR软件 | A cross-platform software for text translation and recognition.
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.