LightX2V
ModelTC
Lightweight Image Video Action Generation Inference Framework
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
33–48 / 351
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 351
ModelTC
Lightweight Image Video Action Generation Inference Framework
FireRedTeam
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering ou…
shivammehta25
[ICASSP 2024] 🍵 Matcha-TTS: A fast TTS architecture with conditional flow matching
R3gm
Synchronized Translation for Videos. Video dubbing
coqui-ai
🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.
kalliope-project
Kalliope is a framework that will help you to create your own personal assistant.
varunshenoy
An extensible, easy-to-use, and portable diffusion web UI 👨🎨
ictnlp
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
cure-lab
[ICLR24] Official implementation of the paper “MagicDrive: Street View Generation with Diverse 3D Geometry Control”
Picovoice
On-device streaming speech-to-text engine powered by deep learning
RVC-Boss
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
FireRedTeam
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switchin…
deepgram
Official Python SDK for Deepgram.
babysor
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
unslothai
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.