No visual example yet
Explore the skillRNSkill Audio to Subtitles
Pluviobyte
Generate real-timestamp subtitle artifacts from final narration audio or merged video with a caption quality gate.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
17–32 / 64
Results: 64
No visual example yet
Explore the skillPluviobyte
Generate real-timestamp subtitle artifacts from final narration audio or merged video with a caption quality gate.
No visual example yet
Explore the skillhuggingface
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and traini…
No visual example yet
Explore the skillabus-aikorea
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTu…
No visual example yet
Explore the skillopen-mmlab
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get…
No visual example yet
Explore the skillspotify
No visual example yet
Explore the skillspotify
A lightweight yet powerful audio-to-MIDI converter with pitch bend detection
No visual example yet
Explore the skillrsxdalv
A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Au…
No visual example yet
Explore the skillmilvus-io
Dealing with all unstructured data, such as reverse image search, audio search, molecular search, video analysis, question and answer systems, NLP, etc.
No visual example yet
Explore the skillCPJKU
Python audio and music signal processing library
No visual example yet
Explore the skillpluja
Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!
No visual example yet
Explore the skillenhuiz
An unofficial PyTorch implementation of the audio LM VALL-E
No visual example yet
Explore the skillreadbeyond
aeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced alignment)
No visual example yet
Explore the skillDamRsn
Audio Plugin for Audio to MIDI transcription using deep learning.
No visual example yet
Explore the skillMiteshPuthran
The neural network model is capable of detecting five different male/female emotions from audio speeches. (Deep Learning, NLP, Python)
No visual example yet
Explore the skilljishengpeng
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
No visual example yet
Explore the skillmravanelli
SincNet is a neural architecture for efficiently processing raw audio samples.