暂未收录效果图
查看技能说明Mlx Audio
Blaizzy
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
OPENAGENTSKILL / DIRECTORY
为下一项任务找到合适的技能。探索适用于 Codex、Claude Code、Cursor 等 Agent 的工具。
46 Skills
搜索结果: 46
暂未收录效果图
查看技能说明Blaizzy
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
暂未收录效果图
查看技能说明pytorch
Data manipulation and transformation for audio signal processing, powered by PyTorch
暂未收录效果图
查看技能说明libAudioFlux
暂未收录效果图
查看技能说明tyiannak
Python Audio Analysis Library: Feature Extraction, Classification, Segmentation and Applications
暂未收录效果图
查看技能说明iver56
A Python library for audio data augmentation. Useful for making audio ML models work well in the real world, not just in the lab.
暂未收录效果图
查看技能说明hkchengrex
[CVPR 2025] MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
暂未收录效果图
查看技能说明diodiogod
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio…
暂未收录效果图
查看技能说明FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
暂未收录效果图
查看技能说明nexu-io
Audio generation skill — jingles, beds, voiceover, and sound effects. Routes music requests to Suno V5 / Udio / Lyria, speech to MiniMax TTS / FishAudio / ElevenLabs V3,…
暂未收录效果图
查看技能说明Pluviobyte
Generate real-timestamp subtitle artifacts from final narration audio or merged video with a caption quality gate.
暂未收录效果图
查看技能说明huggingface
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and traini…
暂未收录效果图
查看技能说明abus-aikorea
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTu…
暂未收录效果图
查看技能说明open-mmlab
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get…
暂未收录效果图
查看技能说明spotify
暂未收录效果图
查看技能说明spotify
暂未收录效果图
查看技能说明rsxdalv
A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Au…