Audio Webui
gitmylo
A webui for different audio related Neural Networks
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
321–336 / 466
Results: 466
gitmylo
A webui for different audio related Neural Networks
mravanelli
SincNet is a neural architecture for efficiently processing raw audio samples.
vargHQ
AI video generation SDK — JSX for videos. One API for Kling, Flux, ElevenLabs, Veed. Built on Vercel AI SDK.
bawangxx
Free and open source text-to-speech software
alan-ai
The Self-Coding System for Your App — Alan AI SDK for Cordova
PowerBeef
Vocello — a local, private voice studio for Apple Silicon. Write a script, pick or describe a voice, and generate speech on-device. macOS today, iPhone coming soon. (For…
alumae
Real-time full-duplex speech recognition server, based on the Kaldi toolkit and the GStreamer framwork.
savbell
💬📝 A small dictation app using OpenAI's Whisper speech recognition model.
SYuan03
Any source (PDF, video, web, audio, text) to interactive learning package with quizzes, flashcards and spaced repetition. One command, 12-section study guide.
yuga-hashimoto
OpenClaw voice assistant app for Android - Wake word activation & system assistant integration
schmitech
A self-hosted AI infrastructure for private RAG and multi-model applications.
antgroup
[AAAI 2026] EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation
PunithVT
🎭 AI Avatar / digital human platform — upload a photo, clone a voice, talk to any face in real time with lip-sync video. Open-source, self-hosted. Claude · Whisper · Ch…
Anil-matcha
🎬 Turn any topic into a finished Vox-style paper-collage explainer / motion graphics video — script, collage keyframes, animation, voice-over, music & captions, all aut…
TensorSpeech
:zap: TensorFlowASR: Almost State-of-the-art Automatic Speech Recognition in Tensorflow 2. Supported languages that can use characters or subwords
netease-youdao
Confucius4-TTS: a Multilingual and Cross-Lingual Zero-Shot TTS Engine
Design direction, Figma implementation, UI review, React performance, browser QA, and safe preview deployments.
Code review, repo inspection, testing, planning, shipping, and engineering workflows for coding agents.
Research, recent-events briefings, PDF parsing, markdown conversion, RAG ingestion, and knowledge workflows.
Stock analysis, market research, quant backtesting, financial data, and investment research skills.
Crawling, scraping, extraction, browser automation, structured data capture, and website-to-markdown workflows.
Presentation generation, editable PPTX, slide decks, speaker notes, and visual storytelling workflows.
Video-generation prompts, B-roll, Vox-style explainers, camera direction, captions, and creative-production workflows.
Image, video, creative production, UI design, multimodal generation, and visual workflow skills.