15 papers
ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
Xiaozhong Lyu, Gen Li, Zhiyin Qian +3
Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environm…
Modeling Subjective Urban Perception with Human Gaze
Lin Che, Xi Wang, Marc Pollefeys +3
Urban perception describes how people subjectively evaluate urban environments, shaping how cities are experienced and understood. Existing computational approaches primarily model…
ScreenParse: Moving Beyond Sparse Grounding with Complete Screen Parsing Supervision
A. Said Gurbuz, Sunghwan Hong, Ahmed Nassar +2
Modern computer-use agents (CUA) must perceive a screen as a structured state, what elements are visible, where they are, and what text they contain, before they can reliably groun…
Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models
Kevin Qu, Haozhe Qi, Mihai Dusmanu +3
Multimodal Large Language Models (MLLMs) have made impressive progress in connecting vision and language, but they still struggle with spatial understanding and viewpoint-aware rea…
LSA: Localized Semantic Alignment for Enhancing Temporal Consistency in Traffic Video Generation
Mirlan Karimov, Teodora Spasojevic, Markus Braun +3
Controllable video generation has emerged as a versatile tool for autonomous driving, enabling realistic synthesis of traffic scenarios. However, existing methods depend on control…
ImLoc: Revisiting Visual Localization with Image-based Representation
Xudong Jiang, Fangjinhua Wang, Silvano Galliani +2
Existing visual localization methods are typically either 2D image-based, which are easy to build and maintain but limited in effective geometric reasoning, or 3D structure-based,…