WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Directorio de skills
Descubre skills reutilizables para AI agents.
Cada recomendación conserva un vínculo claro con su repositorio, auditoría y ruta de instalación.
Resultados de búsqueda: voice-recognition
Directorio en inglés1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
The world's simplest facial recognition api for Python and the command line
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
Rewrites AI-sounding text so it reads naturally while preserving every factual claim and matching the writer's voice.
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
A Lightweight Face Recognition and Facial Attribute Analysis (Age, Gender, Emotion and Race) Library for Python
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
💬 Open source machine learning framework to automate text- and voice-based conversations: NLU, dialogue management, connect to Slack, Facebook, and more - Create chatbots and voice assistants
🌈一个跨平台的划词翻译和OCR软件 | A cross-platform software for text translation and recognition.
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.