4 papers
UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
Zhipeng Bao, Zhen Zhu, Nupur Kumari +4
Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handled by separate models. We study…
Walk through Paintings: Egocentric World Models from Internet Priors
Anurag Bagchi, Zhipeng Bao, Homanga Bharadhwaj +3
What if a video generation model could not only imagine a plausible future, but the correct one -- accurately reflecting how the world changes with each action? We answer this by p…
ReferEverything: Towards Segmenting Everything We Can Speak of in Videos
Anurag Bagchi, Zhipeng Bao, Yu-Xiong Wang +2
We present REM, a framework for segmenting a wide range of concepts in video that can be described through natural language. Our method leverages the universal visual-language mapp…
Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
Yunze Man, Shuhong Zheng, Zhipeng Bao +3
Complex 3D scene understanding has gained increasing attention, with scene encoding strategies playing a crucial role in this success. However, the optimal scene encoding strategie…