3 papers
cs.CV2026
UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
Zhipeng Bao, Zhen Zhu, Nupur Kumari +4
Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handled by separate models. We study…
cs.CV2026
Walk through Paintings: Egocentric World Models from Internet Priors
Anurag Bagchi, Zhipeng Bao, Homanga Bharadhwaj +3
What if a video generation model could not only imagine a plausible future, but the correct one -- accurately reflecting how the world changes with each action? We answer this by p…
cs.CV2025
ReferEverything: Towards Segmenting Everything We Can Speak of in Videos
Anurag Bagchi, Zhipeng Bao, Yu-Xiong Wang +2
We present REM, a framework for segmenting a wide range of concepts in video that can be described through natural language. Our method leverages the universal visual-language mapp…