StableAvatar
Francis-Rings
We present StableAvatar, the first end-to-end video diffusion transformer, which synthesizes infinite-length high-quality audio-driven avatar videos without any post-pro…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
49–64 / 122
Results: 122
Francis-Rings
We present StableAvatar, the first end-to-end video diffusion transformer, which synthesizes infinite-length high-quality audio-driven avatar videos without any post-pro…
ekwek1
Soprano: Instant, Ultra-Realistic Text-to-Speech
avevlad
Список русскоязычных подкастов на тему информационных технологий
nari-labs
TTS model capable of streaming conversational audio in realtime.
huggingface
Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.
alphacep
Offline speech recognition for Android with Vosk library.
audeering
Python package for openSMILE
chenyme
这是一个全自动(音频)视频翻译项目。利用Whisper识别声音,AI大模型翻译字幕,最后合并字幕视频,生成翻译后的视频。
keithito
A TensorFlow implementation of Google's Tacotron speech synthesis with pre-trained model (unofficial)
enhuiz
An unofficial PyTorch implementation of the audio LM VALL-E
mybigday
React Native binding of whisper.cpp.
Camb-ai
MARS5 speech model (TTS) from CAMB.AI
DamRsn
Audio Plugin for Audio to MIDI transcription using deep learning.
marytts
MARY TTS -- an open-source, multilingual text-to-speech synthesis system written in pure java
asticode
Golang framework to build an AI that can understand and speak back to you, and everything else you want
jamsch
Speech Recognition for React Native Expo projects
Design direction, Figma implementation, UI review, React performance, browser QA, and safe preview deployments.
Code review, repo inspection, testing, planning, shipping, and engineering workflows for coding agents.
Research, recent-events briefings, PDF parsing, markdown conversion, RAG ingestion, and knowledge workflows.
Stock analysis, market research, quant backtesting, financial data, and investment research skills.
Crawling, scraping, extraction, browser automation, structured data capture, and website-to-markdown workflows.
Presentation generation, editable PPTX, slide decks, speaker notes, and visual storytelling workflows.
Video-generation prompts, B-roll, Vox-style explainers, camera direction, captions, and creative-production workflows.
Image, video, creative production, UI design, multimodal generation, and visual workflow skills.