Infinity
FoundationVision
[CVPR 2025 Oral]Infinity ∞ : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
161–176 / 478
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 478
FoundationVision
[CVPR 2025 Oral]Infinity ∞ : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
OpenGVLab
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, S…
jy0205
[ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling
tnfe
A fast video processing library based on node.js (一个基于node.js的高速视频制作库)
lifeiteng
PyTorch implementation of VALL-E(Zero-Shot Text-To-Speech), Reproduced Demo https://lifeiteng.github.io/valle/index.html
DigitalPhonetics
Controllable and fast Text-to-Speech for over 7000 languages!
pluja
Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!
diodiogod
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio…
Doubiiu
[ECCV 2024, Oral] DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors
williamyang1991
[SIGGRAPH Asia 2023] Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation
byjlw
Analyze videos using LLMs, Computer Vision and Automatic Speech Recognition
readbeyond
aeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced alignment)
TMElyralab
MuseV: Infinite-length and High Fidelity Virtual Human Video Generation with Visual Conditioned Parallel Denoising