Skill-Verzeichnis

Wiederverwendbare Skills für AI Agents entdecken.

Durchsuche reale GitHub-Skills nach Aufgabe und prüfe Stars, Trust, Audit, Kategorie und Installationspfad vor der Verwendung.

Jede Empfehlung bleibt mit ihrem Repository, Audit und Installationspfad nachvollziehbar.

Suchergebnisse: voice-recognition

Englisches Verzeichnis

WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

23K
Stars
87/100
Trust
Kategorie: media-automationAudit

1 min voice data can also be used to train a good TTS model! (few shot voice cloning)

59K
Stars
87/100
Trust
Kategorie: media-automationAudit

The world's simplest facial recognition api for Python and the command line

57K
Stars
77/100
Trust
Kategorie: ml-automationAudit

Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node

15K
Stars
87/100
Trust
Kategorie: media-automationAudit

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

31K
Stars
87/100
Trust
Kategorie: media-automationAudit

Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.

8.6K
Stars
77/100
Trust
Kategorie: media-automationAudit

A Lightweight Face Recognition and Facial Attribute Analysis (Age, Gender, Emotion and Race) Library for Python

23K
Stars
87/100
Trust
Kategorie: ml-automationAudit

Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

22K
Stars
87/100
Trust
Kategorie: support-automationAudit

💬 Open source machine learning framework to automate text- and voice-based conversations: NLU, dialogue management, connect to Slack, Facebook, and more - Create chatbots and voice assistants

21K
Stars
86/100
Trust
Kategorie: ml-automationAudit

🌈一个跨平台的划词翻译和OCR软件 | A cross-platform software for text translation and recognition.

19K
Stars
87/100
Trust
Kategorie: document-processingAudit

Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.

1.5K
Stars
74/100
Trust
Kategorie: design-creativeAudit

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

12K
Stars
87/100
Trust
Kategorie: media-automationAudit