6 papers
When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
Yu Fang, Yuchun Feng, Dong Jing +5
The paper studies how Vision-Language-Action (VLA) models often ignore language instructions by relying on visual shortcuts, introduces a counterfactual benchmark (LIBERO-CF) to ev…
DenseReward: Dense Reward Learning via Failure Synthesis for Robotic Manipulation
Yu Fang, Wanxi Dong, Jiaqi Liu +7
The paper presents DenseReward, a dense visual‑language reward model for robotic manipulation that is trained on automatically synthesized failure trajectories in simulation, enabl…
Current as Touch: Proprioceptive Contact Feedback for Compliant Dexterous Manipulation
Chenyang Ma, Yunchao Yao, Zhenyu Wei +3
Compliance is essential for dexterous manipulation, yet existing solutions often rely on external tactile or force sensors that are costly, fragile, and difficult to deploy on low-…
Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
Yu Fang, Kanchana Ranasinghe, Le Xue +10
Vision-Language-Action (VLA) models have achieved remarkable progress in robotic manipulation by mapping multimodal observations and instructions directly to actions. However, they…
ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis
Yu Fang, Yue Yang, Xinghao Zhu +4
Vision-language-action (VLA) models present a promising paradigm by training policies directly on real robot datasets like Open X-Embodiment. However, the high cost of real-world d…
BOSS: Benchmark for Observation Space Shift in Long-Horizon Task
Yue Yang, Linfeng Zhao, Mingyu Ding +2
Robotics has long sought to develop visual-servoing robots capable of completing previously unseen long-horizon tasks. Hierarchical approaches offer a pathway for achieving this go…