40 papers
MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
Wenjie Zhu, Yabin Zhang, Wenjun Zeng +1
Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable M…
GEM-Occ: From Visual Geometry Evidence to Embodied Semantic Occupancy Memory
Hu Zhu, Bohan Li, Xianda Guo +6
Semantic occupancy provides a structured spatial memory for embodied indoor agents by jointly representing occupied regions, observed free space, unknown areas, and object semantic…
Bridging 3D Gaussians and Semantic Occupancy for Comprehensive Open-Vocabulary Scene Understanding from Unposed Images
Hu Zhu, Bohan Li, Xianda Guo +5
Comprehensive 3D scene understanding from sparse, unposed images requires a model to recover renderable geometry, open-vocabulary semantics, and free/occupied 3D space without rely…
OmniNWM: Omniscient Driving Navigation World Models
Bohan Li, Zhuang Ma, Dalong Du +10
Autonomous driving world models are expected to work effectively across three core dimensions: state, action, and reward. However, existing methods are typically restricted to frag…
DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning
Zili Lin, Wenyao Zhang, Yuyang Zhang +9
Demonstration augmentation is proposed for cost-efficient data acquisition, but existing methods are fundamentally limited in deformable manipulation due to two challenges: (1) the…
Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs
Wenjie Zhu, Yabin Zhang, Liang Xu +3
While test-time adaptation (TTA) empowers vision-language models to adapt without costly retraining, it remains highly vulnerable to out-of-distribution (OOD) outliers prevalent in…