Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
PROSPECT: Unified Streaming Vision-Language Navigation via Semantic--Spatial Fusion and Latent Predictive Representation
Zehua Fan, Wenqi Lyu, Wenxuan Song +12
Multimodal large language models (MLLMs) have advanced zero-shot end-to-end Vision-Language Navigation (VLN), yet robust navigation requires not only semantic understanding but als…
cs.CV2025
VisualChef: Generating Visual Aids in Cooking via Mask Inpainting
Oleh Kuzyk, Zuoyue Li, Marc Pollefeys +1
Cooking requires not only following instructions but also understanding, executing, and monitoring each step - a process that can be challenging without visual guidance. Although r…
cs.CV2024
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
Mariam Hassan, Sebastian Stapf, Ahmad Rahimi +17
We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, ou…