No visual example yet
Explore the skillSenseVoice
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
OPENAGENTSKILL / DIRECTORY
Finde den passenden Skill für deine nächste Aufgabe mit Codex, Claude Code, Cursor und mehr.
64 Skills
Ergebnisse: 64
No visual example yet
Explore the skillFunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
No visual example yet
Explore the skillqxresearch
Python hands on tutorial with 50+ Python Application (10 lines of code) By @xiaowuc2
No visual example yet
Explore the skillBlaizzy
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
No visual example yet
Explore the skillpytorch
Data manipulation and transformation for audio signal processing, powered by PyTorch
No visual example yet
Explore the skilllibAudioFlux
A library for audio and music analysis, feature extraction.
No visual example yet
Explore the skilliver56
A Python library for audio data augmentation. Useful for making audio ML models work well in the real world, not just in the lab.
No visual example yet
Explore the skillhkchengrex
[CVPR 2025] MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
No visual example yet
Explore the skilldiodiogod
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio…
No visual example yet
Explore the skillmilvus-io
Dealing with all unstructured data, such as reverse image search, audio search, molecular search, video analysis, question and answer systems, NLP, etc.
No visual example yet
Explore the skillnexu-io
Audio generation skill — jingles, beds, voiceover, and sound effects. Routes music requests to Suno V5 / Udio / Lyria, speech to MiniMax TTS / FishAudio / ElevenLabs V3,…
No visual example yet
Explore the skillPluviobyte
Generate real-timestamp subtitle artifacts from final narration audio or merged video with a caption quality gate.
No visual example yet
Explore the skillEventual-Inc
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
No visual example yet
Explore the skillargoproj
Event-driven Automation Framework for Kubernetes
No visual example yet
Explore the skillknative
Event-driven application platform for Kubernetes
No visual example yet
Explore the skillhuggingface
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and traini…
No visual example yet
Explore the skilljacobgil
Advanced AI Explainability for computer vision. Support for CNNs, Vision Transformers, Classification, Object detection, Segmentation, Image similarity and more.