Torchscale
microsoft
Foundation Architecture for (M)LLMs
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
241–256 / 466
Results: 466
microsoft
Foundation Architecture for (M)LLMs
AutoArk
[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!
chenyme
这是一个全自动(音频)视频翻译项目。利用Whisper识别声音,AI大模型翻译字幕,最后合并字幕视频,生成翻译后的视频。
pluja
Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!
keithito
A TensorFlow implementation of Google's Tacotron speech synthesis with pre-trained model (unofficial)
enhuiz
An unofficial PyTorch implementation of the audio LM VALL-E
tejaswigowda
A browser-based video editor powered by ffmpeg.wasm. No uploads, no servers -- all processing happens locally in your browser using WebAssembly.
AliAkhtari78
Extract public Spotify data — tracks, albums, artists, playlists, podcasts & lyrics — without the official API. Sync + async, typed models, one dependency.
mybigday
React Native binding of whisper.cpp.
VRCWizard
Speech to Text to Speech. Song now playing. Sends text as OSC messages to VRChat to display on avatar. (STTTS) (Speech to TTS) (VRC STT System) (VTuber TTS)
readbeyond
aeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced alignment)
zzw922cn
End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow
Camb-ai
MARS5 speech model (TTS) from CAMB.AI
DamRsn
Audio Plugin for Audio to MIDI transcription using deep learning.
hahahumble
💬 SpeechGPT is a web application that enables you to converse with ChatGPT.
cboard-org
Augmentative and Alternative Communication (AAC) system with text-to-speech for the browser
Design direction, Figma implementation, UI review, React performance, browser QA, and safe preview deployments.
Code review, repo inspection, testing, planning, shipping, and engineering workflows for coding agents.
Research, recent-events briefings, PDF parsing, markdown conversion, RAG ingestion, and knowledge workflows.
Stock analysis, market research, quant backtesting, financial data, and investment research skills.
Crawling, scraping, extraction, browser automation, structured data capture, and website-to-markdown workflows.
Presentation generation, editable PPTX, slide decks, speaker notes, and visual storytelling workflows.
Video-generation prompts, B-roll, Vox-style explainers, camera direction, captions, and creative-production workflows.
Image, video, creative production, UI design, multimodal generation, and visual workflow skills.