Dreamtalk
ali-vilab
Official implementations for paper: DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic Models
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
33–48 / 56
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 56
ali-vilab
Official implementations for paper: DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic Models
Tencent-Hunyuan
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
VideoFlint
A video composition framework build on top of AVFoundation. It's simple to use and easy to extend.
Saik0s
The open-source iOS app that's making quality voice transcription more accessible on mobile devices.
MiteshPuthran
The neural network model is capable of detecting five different male/female emotions from audio speeches. (Deep Learning, NLP, Python)
jishengpeng
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
ictnlp
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
mravanelli
SincNet is a neural architecture for efficiently processing raw audio samples.
tejaswigowda
A browser-based video editor powered by ffmpeg.wasm. No uploads, no servers -- all processing happens locally in your browser using WebAssembly.
antgroup
[AAAI 2026] EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation
fluxions-ai
Real-time voice assistant — WebRTC streaming, faster-whisper ASR, local LLM, Vui Nano (300M) TTS. OpenAI Realtime API compatible. Voice cloning, barge-in, ~9× realtime o…
rapidaai
Rapida is an open-source, end-to-end voice AI orchestration platform for building real-time conversational voice agents with audio streaming, STT, TTS, VAD, multi-channe…
aetaric
Checkrr Scans your library files for corrupt media and optionally replaces the files via sonarr and radarr
FireRedTeam
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switchin…
izwi-ai
Voice AI runtime. Local first transcription, speaker diarization, TTS, and voice cloning with an OpenAI compatible API.
huggingface
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.