No visual example yet
Explore the skillVoice Pro
abus-aikorea
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTu…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
17–32 / 55
Results: 55
No visual example yet
Explore the skillabus-aikorea
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTu…
No visual example yet
Explore the skillopen-mmlab
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get…
No visual example yet
Explore the skillspotify
No visual example yet
Explore the skillspotify
A lightweight yet powerful audio-to-MIDI converter with pitch bend detection
No visual example yet
Explore the skillrsxdalv
A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Au…
No visual example yet
Explore the skillmilvus-io
Dealing with all unstructured data, such as reverse image search, audio search, molecular search, video analysis, question and answer systems, NLP, etc.
No visual example yet
Explore the skillCPJKU
Python audio and music signal processing library
No visual example yet
Explore the skillpluja
Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!
No visual example yet
Explore the skillenhuiz
An unofficial PyTorch implementation of the audio LM VALL-E
No visual example yet
Explore the skillreadbeyond
aeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced alignment)
No visual example yet
Explore the skillDamRsn
Audio Plugin for Audio to MIDI transcription using deep learning.
No visual example yet
Explore the skillMiteshPuthran
The neural network model is capable of detecting five different male/female emotions from audio speeches. (Deep Learning, NLP, Python)
No visual example yet
Explore the skilljishengpeng
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
No visual example yet
Explore the skillmravanelli
SincNet is a neural architecture for efficiently processing raw audio samples.
No visual example yet
Explore the skillPluviobyte
Generate a controlled local narration workflow with auditions, version tracking, and subtitle-ready final audio.
No visual example yet
Explore the skillbackblaze-labs
Genblaze is an open source Python SDK for orchestrating generative AI media pipelines across video, audio, and image providers with built in provenance for every output.