Whisper Playground
saharmor
Build real time speech2text web apps using OpenAI's Whisper https://openai.com/blog/whisper/
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
Page 25 · 16 shown · 744 public entries
Results: 744
saharmor
Build real time speech2text web apps using OpenAI's Whisper https://openai.com/blog/whisper/
thedivergentai
Expert patterns for AnimationPlayer including track types (Value, Method, Audio, Bezier), root motion extraction, animation callbacks, procedural animation generation, c…
thedivergentai
Expert patterns for Godot AutoLoad (singleton) architecture including global state management, scene transitions, signal-based communication, dependency injection, autol…
BandarLabs
Convert any git repository into an engaging podcast
staruhub
用火山引擎 Podcast AI 模型生成中文双人对话播客。当用户要把文章、报告、话题文本转成播客音频、生成对话式音频内容时使用,需要环境具备火山引擎 APP_ID 和 ACCESS_KEY。支持 mp3/ogg_opus/pcm/aac、语速调节、自定义音色、断点续传。不用于:单人朗读式 TTS(用普通语音合成)、英文播客(模型主要优…
preziotte
An experimental music visualizer using d3.js and the web audio api.
gexgd0419
Make Azure natural TTS voices accessible to any SAPI 5-compatible application.
rnchg
AI Productivity Tool - Free and open source, improve user productivity, and protect privacy and data security. Including but not limited to: built-in local exclusive Cha…
kishanrajput23
A python based desktop voice assistant capable of executing system-level commands, integrating speech recognition and text-to-speech, and handling asynchronous user inte…
yeyupiaoling
基于PaddlePaddle实现的语音识别,中文语音识别。项目完善,识别效果好。支持Windows,Linux下训练和预测,支持Nvidia Jetson开发板预测。
HKUDS
"VideoAgent: All-in-One Agentic Framework for Video Understanding, Editing, and Remaking"
berabuddies
Native video analysis using Google Gemini API. Upload and analyze video files — describe scenes, extract text/UI, answer questions about content, transcribe speech, iden…
markovka17
Deep learning for audio processing
jame581
Use when importing and managing assets — image compression, 3D scene import, audio formats, resource formats, and import configuration
dlutton
amanvirparhar
A real-time silent speech recognition tool.