Phantom
Phantom-video
Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
209–224 / 478
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 478
Phantom-video
Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
Azure-Samples
Microsoft Text-to-Speech API sample code in several languages, part of Cognitive Services.
thu-ml
[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.
superstreamerapp
An open, scalable, online streaming setup. All-in-one toolkit from ingest to adaptive video playback. Built for developers in need of video tooling.
FoundationVision
Autoregressive Model Beats Diffusion: 🦙 Llama for Scalable Image Generation
travisvn
Free, high-quality text-to-speech API endpoint to replace OpenAI, Azure, or ElevenLabs
julius-speech
Open-Source Large Vocabulary Continuous Speech Recognition Engine
bytedance
[ICCV 2025] 🔥🔥 UNO: A Universal Customization Method for Both Single and Multi-Subject Conditioning
Softcatala
Whisper command line client compatible with original OpenAI client based on CTranslate2.
syhw
Attempt at tracking states of the arts and recent results (bibliography) on speech recognition.
Francis-Rings
We present StableAvatar, the first end-to-end video diffusion transformer, which synthesizes infinite-length high-quality audio-driven avatar videos without any post-pro…
ekwek1
Soprano: Instant, Ultra-Realistic Text-to-Speech
harskish
Discovering Interpretable GAN Controls [NeurIPS 2020]
ali-vilab
Official implementations for paper: DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic Models
kalliope-project
Kalliope is a framework that will help you to create your own personal assistant.