Skill-Verzeichnis

Wiederverwendbare Skills für AI Agents entdecken.

Durchsuche reale GitHub-Skills nach Aufgabe und prüfe Stars, Trust, Audit, Kategorie und Installationspfad vor der Verwendung.

Jede Empfehlung bleibt mit ihrem Repository, Audit und Installationspfad nachvollziehbar.

Suchergebnisse: speech-translation

Englisches Verzeichnis

WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

23K
Stars
87/100
Trust
Kategorie: media-automationAudit

Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node

15K
Stars
87/100
Trust
Kategorie: media-automationAudit

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

31K
Stars
87/100
Trust
Kategorie: media-automationAudit

Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.

8.6K
Stars
77/100
Trust
Kategorie: media-automationAudit

An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

21K
Stars
77/100
Trust
Kategorie: media-automationAudit

🌈一个跨平台的划词翻译和OCR软件 | A cross-platform software for text translation and recognition.

19K
Stars
87/100
Trust
Kategorie: document-processingAudit

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages

13K
Stars
87/100
Trust
Kategorie: media-automationAudit

Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.

13K
Stars
87/100
Trust
Kategorie: media-automationAudit

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

12K
Stars
87/100
Trust
Kategorie: media-automationAudit

A PyTorch-based Speech Toolkit

12K
Stars
82/100
Trust
Kategorie: media-automationAudit

A generative speech model for daily dialogue.

40K
Stars
81/100
Trust
Kategorie: agent-frameworksAudit

🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time

37K
Stars
76/100
Trust
Kategorie: media-automationAudit