6 papers
Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies
Fangyuan Wang, Peng Zhou, Jiaming Qi +4
Vision-language-action (VLA) models typically inject proprioception only as a late conditioning signal, preventing robot state from grounding instruction understanding or directing…
REF-VLM: Triplet-Based Referring Paradigm for Unified Visual Decoding
Yan Tai, Luhao Zhu, Yunan Ding +4
Multimodal Large Language Models (MLLMs) demonstrate robust zero-shot capabilities across diverse vision-language tasks after training on mega-scale datasets. However, dense predic…
DiffRefiner: Coarse to Fine Trajectory Planning via Diffusion Refinement with Semantic Interaction for End to End Autonomous Driving
Liuhan Yin, Runkun Ju, Guodong Guo +1
Unlike discriminative approaches in autonomous driving that predict a fixed set of candidate trajectories of the ego vehicle, generative methods, such as diffusion models, learn th…
Phy-Tac: Toward Human-Like Grasping via Physics-Conditioned Tactile Goals
Shipeng Lyu, Lijie Sheng, Fangyuan Wang +5
Humans naturally grasp objects with minimal level required force for stability, whereas robots often rely on rigid, over-squeezing control. To narrow this gap, we propose a human-i…
HuBE: Cross-Embodiment Human-like Behavior Execution for Humanoid Robots
Shipeng Lyu, Fangyuan Wang, Weiwei Lin +3
Achieving both behavioral similarity and appropriateness in human-like motion generation for humanoid robot remains an open challenge, further compounded by the lack of cross-embod…
Instruction-Augmented Long-Horizon Planning: Embedding Grounding Mechanisms in Embodied Mobile Manipulation
Fangyuan Wang, Shipeng Lyu, Peng Zhou +3
Enabling humanoid robots to perform long-horizon mobile manipulation planning in real-world environments based on embodied perception and comprehension abilities has been a longsta…