Vall E
lifeiteng
PyTorch implementation of VALL-E(Zero-Shot Text-To-Speech), Reproduced Demo https://lifeiteng.github.io/valle/index.html
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
65–80 / 251
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 251
lifeiteng
PyTorch implementation of VALL-E(Zero-Shot Text-To-Speech), Reproduced Demo https://lifeiteng.github.io/valle/index.html
Doubiiu
[ECCV 2024, Oral] DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors
byjlw
Analyze videos using LLMs, Computer Vision and Automatic Speech Recognition
TMElyralab
MuseV: Infinite-length and High Fidelity Virtual Human Video Generation with Visual Conditioned Parallel Denoising
deepgram
Official JavaScript SDK for Deepgram.
FireRedTeam
FireRed-Image-Edit is a powerful image editing foundation model achieving open-source state-of-the-art performance with precise instruction following, high-fidelity gene…
Rayhane-mamah
DeepMind's Tacotron-2 Tensorflow implementation
pannous
🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks
Phantom-video
Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
FoundationVision
Autoregressive Model Beats Diffusion: 🦙 Llama for Scalable Image Generation
Softcatala
Whisper command line client compatible with original OpenAI client based on CTranslate2.
Francis-Rings
We present StableAvatar, the first end-to-end video diffusion transformer, which synthesizes infinite-length high-quality audio-driven avatar videos without any post-pro…