6 papers
Latent Visual States for Efficient Multimodal Reasoning
Xiuwei Chen, Wentao Hu, Yongxin Wang +8
The integration of visual evidence has significantly enhanced the capabilities of large multimodal models. However, this integration predominantly relies on generating discrete out…
Intend, Reflect, Refine: An Adaptive Multimodal Reflection Framework for Autonomous Driving
Zisheng Chen, Yuping Qiu, Jianhua Han +6
Recent Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving by incorporating reasoning for better interpretability and planning quality. However, most ex…
SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning
Kun Xiang, Terry Jingchen Zhang, Zirong Liu +15
We introduce SeePhys Pro, a fine-grained modality transfer benchmark that studies whether models preserve the same reasoning capability when critical information is progressively t…
E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion
Zhihao Zhan, Jiaying Zhou, Likui Zhang +10
Vision-Language-Action (VLA) models offer a unified framework for robotic manipulation by integrating visual perception, language understanding, and control generation. However, ex…
AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots
Likui Zhang, Tao Tang, Zhihao Zhan +9
Recent advances in Visual-Language-Action (VLA) models have shown promising potential for robotic manipulation tasks. However, real-world robotic tasks often involve long-horizon,…
RoboPearls: Editable Video Simulation for Robot Manipulation
Tao Tang, Likui Zhang, Youpeng Wen +9
The development of generalist robot manipulation policies has seen significant progress, driven by large-scale demonstration data across diverse environments. However, the high cos…