Generate and monitor CyberBara Public API v1 image, video, audio, and music tasks end-to-end. Use when work involves CyberBara `/api/v1` endpoints for listing models, up…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
Page 37 · 16 shown · 743 public entries
Results: 743
Generate and monitor CyberBara Public API v1 image, video, audio, and music tasks end-to-end. Use when work involves CyberBara `/api/v1` endpoints for listing models, up…
jackwuwei
The ChatGPT/DeepSeek Voice Assistant uses a Raspberry Pi (or desktop) to enable spoken conversation with OpenAI or DeepSeek large language models. This implementation li…
wizgrav
Application of music theory in audio reactive visualizations
X-LANCE
[ICASSP 2024] This is the official code for "VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching"
hassancs91
Final step of the AI Storybook pipeline. Consolidates the scenes, images, and per-scene audio into ONE self-contained HTML storybook — a swipe/tap player with every imag…
mturac
Unified media generation via fal.ai MCP — image, video, and audio. Covers text-to-image (Nano Banana), text/image-to-video (Seedance, Kling, Veo 3), text-to-speech (CSM-…
scopeInfinity
Video to Text: Natural language description generator for some given video. [Video Captioning]
Nikorasu
A nearly-live implementation of OpenAI's Whisper, using sounddevice. Requires existing Whisper install.
deterministic-algorithms-lab
Tacotron 2 - PyTorch implementation with faster-than-realtime inference modified to enable cross lingual voice cloning.
megaease
EaseVoice Trainer is a simple and user-friendly voice cloning and speech model trainer.
EtienneAb3d
Experimental code: sound file preprocessing to optimize Whisper transcriptions without hallucinated texts
keonlee9420
PyTorch Implementation of DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs
dmstern
[DEPRECATED] A browser extension to block likers, retweeters, list members and Twitter ads and share your block lists with others. - say NO to hate speech!
jingweizhanghuai
Morn是一个C语言的基础工具和基础算法库,包括数据结构、图像处理、音频处理、机器学习等,具有简单、通用、高效的特点。
rhulha
Unlimited text-to-speech in the Browser using Kokoro-JS, 100% local, 100% open source
keonlee9420
PyTorch Implementation of PortaSpeech: Portable and High-Quality Generative Text-to-Speech