Video Analyzer
byjlw
Analyze videos using LLMs, Computer Vision and Automatic Speech Recognition
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
65–80 / 382
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 382
byjlw
Analyze videos using LLMs, Computer Vision and Automatic Speech Recognition
marytts
MARY TTS -- an open-source, multilingual text-to-speech synthesis system written in pure java
FireRedTeam
FireRed-Image-Edit is a powerful image editing foundation model achieving open-source state-of-the-art performance with precise instruction following, high-fidelity gene…
Rayhane-mamah
DeepMind's Tacotron-2 Tensorflow implementation
pannous
🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks
Phantom-video
Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
julius-speech
Open-Source Large Vocabulary Continuous Speech Recognition Engine
Softcatala
Whisper command line client compatible with original OpenAI client based on CTranslate2.
syhw
Attempt at tracking states of the arts and recent results (bibliography) on speech recognition.
Francis-Rings
We present StableAvatar, the first end-to-end video diffusion transformer, which synthesizes infinite-length high-quality audio-driven avatar videos without any post-pro…
ekwek1
Soprano: Instant, Ultra-Realistic Text-to-Speech
harskish
Discovering Interpretable GAN Controls [NeurIPS 2020]
ali-vilab
Official implementations for paper: DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic Models