Autoregressive Models In Vision Survey
ChaofanTao
[TMLR 2025🔥] A survey for the autoregressive models in vision.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 19
Results: 19
ChaofanTao
[TMLR 2025🔥] A survey for the autoregressive models in vision.
FoundationVision
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation…
FoundationVision
[CVPR 2025 Oral]Infinity ∞ : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
thu-ml
[ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation" & Causal For…
FoundationVision
[NeurIPS 2025 Oral]Infinity⭐️: Unified Spacetime AutoRegressive Modeling for Visual Generation
baaivision
[ICLR 2025] Autoregressive Video Generation without Vector Quantization
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
FoundationVision
Autoregressive Model Beats Diffusion: 🦙 Llama for Scalable Image Generation
opendatalab
A diffusion-based framework for document OCR that replaces autoregressive decoding with block-level parallel diffusion decoding.
lucidrains
Implementation of E2-TTS, "Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS", in Pytorch
baudm
Scene Text Recognition with Permuted Autoregressive Sequence Models (ECCV 2022)
lucidrains
Implementation of Q-Transformer, Scalable Offline Reinforcement Learning via Autoregressive Q-Functions, out of Google Deepmind
keonlee9420
A Non-Autoregressive Transformer based Text-to-Speech, supporting a family of SOTA transformers with supervised and unsupervised duration modelings. This project grows w…
keonlee9420
PyTorch Implementation of Non-autoregressive Expressive (emotional, conversational) TTS based on FastSpeech2, supporting English, Korean, and your own languages.
mbzuai-oryx
LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.