4 papers
PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation
Yutai Li, Shaohui Peng, Jiaming Guo +8
Vision-Language-Action (VLA) models offer a promising paradigm for generalist robotic policies, yet their adaptation is hindered by data inefficiency and poor generalization. We ar…
Run, Ruminate, and Regulate: A Dual-process Thinking System for Vision-and-Language Navigation
Yu Zhong, Zihao Zhang, Rui Zhang +9
Vision-and-Language Navigation (VLN) requires an agent to dynamically explore complex 3D environments following human instructions. Recent research underscores the potential of har…
World-Consistent Data Generation for Vision-and-Language Navigation
Yu Zhong, Rui Zhang, Zihao Zhang +9
Vision-and-Language Navigation (VLN) is a challenging task that requires an agent to navigate through photorealistic environments following natural-language instructions. One main…
Object-Level Verbalized Confidence Calibration in Vision-Language Models via Semantic Perturbation
Yunpu Zhao, Rui Zhang, Junbin Xiao +5
Vision-language models (VLMs) excel in various multimodal tasks but frequently suffer from poor calibration, resulting in misalignment between their verbalized confidence and respo…