activity
20242026
collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking

Jan Kulik, Bjarni Dagur Thor Karason, Yung-Hsu Yang +3

Understanding 3D scenes from egocentric video is fundamental for robotics and autonomous navigation, yet rapid viewpoint changes and partial occlusions make building structured rep…

cs.CV2026

MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight

Zehua Fan, Junjie He, Wenxuan Song +14

World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands sim…

cs.CV2026

Modeling Subjective Urban Perception with Human Gaze

Lin Che, Xi Wang, Marc Pollefeys +3

Urban perception describes how people subjectively evaluate urban environments, shaping how cities are experienced and understood. Existing computational approaches primarily model…

cs.CV2026

PROSPECT: Unified Streaming Vision-Language Navigation via Semantic--Spatial Fusion and Latent Predictive Representation

Zehua Fan, Wenqi Lyu, Wenxuan Song +12

Multimodal large language models (MLLMs) have advanced zero-shot end-to-end Vision-Language Navigation (VLN), yet robust navigation requires not only semantic understanding but als…

cs.CV2025

VisualChef: Generating Visual Aids in Cooking via Mask Inpainting

Oleh Kuzyk, Zuoyue Li, Marc Pollefeys +1

Cooking requires not only following instructions but also understanding, executing, and monitoring each step - a process that can be challenging without visual guidance. Although r…

cs.CV2024

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

Mariam Hassan, Sebastian Stapf, Ahmad Rahimi +17

We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, ou…