9 papers
When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
Yu Fang, Yuchun Feng, Dong Jing +5
The paper studies how Vision-Language-Action (VLA) models often ignore language instructions by relying on visual shortcuts, introduces a counterfactual benchmark (LIBERO-CF) to ev…
DenseReward: Dense Reward Learning via Failure Synthesis for Robotic Manipulation
Yu Fang, Wanxi Dong, Jiaqi Liu +7
The paper presents DenseReward, a dense visual‑language reward model for robotic manipulation that is trained on automatically synthesized failure trajectories in simulation, enabl…
HumanoidMimicGen: Data Generation for Loco-Manipulation via Whole-Body Planning
Kevin Lin, Ajay Mandlekar, Caelan Reed Garrett +7
Imitation learning is a promising approach for training humanoid robots to both walk and manipulate, but it requires a large number of demonstrations, which are time-intensive and…
INHerit-SG: Incremental Hierarchical Semantic Scene Graphs with RAG-Style Retrieval
YukTungSamuel Fang, Zhikang Shi, Jiabin Qiu +5
Driven by recent advancements in foundation models, semantic scene graphs have emerged as a promising paradigm for high-level 3D environmental abstraction in robot navigation. Howe…
DreamGen: Unlocking Generalization in Robot Learning through Video World Models
Joel Jang, Seonghyeon Ye, Zongyu Lin +25
We introduce DreamGen, a simple yet highly effective 4-stage pipeline for training robot policies that generalize across behaviors and environments through neural trajectories - sy…
FLARE: Robot Learning with Implicit World Modeling
Ruijie Zheng, Jing Wang, Scott Reed +18
We introduce uture tent presentation Alignment (), a novel framework that integrates predictive latent world modeling into rob…