Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
Baifeng Shi, Stephanie Fu, Long Lian +10
Multi-modal large language models (MLLMs) have advanced general-purpose video understanding but struggle with long, high-resolution videos -- they process every pixel equally in th…
cs.CV2026
Progressive Checkerboards for Autoregressive Multiscale Image Generation
David Eigen
A key challenge in autoregressive image generation is to efficiently sample independent locations in parallel, while still modeling mutual dependencies with serial conditioning. So…