5 papers
LM-X: Explainable Vision--Language--Action Modeling via Progress, Event, and Uncertainty Prediction
Jin Lou, Zhiyuan Jing, Xupeng Wang +21
Large-scale vision--language--action (VLA) policies have advanced generalist robot control, yet most remain stimulus-to-action black boxes: actions are exposed, but their explanato…
RoboCousin: Build Your Own Simulation Playground for Robust Bimanual Robotic Manipulation
Jingxuan Zhu, Jingyi Li, LiangLiang Chen +3
Bimanual manipulation policies require large and diverse training datasets, yet collecting demonstrations on physical robots is expensive and difficult to scale. Simulation can gen…
AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly
Zhi Jing, Jinbin Qiao, Ouyang Lu +5
Spatial reasoning is a fundamental capability for embodied intelligence, especially for fine-grained manipulation tasks such as robotic assembly. Recent methods based on vision-lan…
HumanoidGen: Data Generation for Bimanual Dexterous Manipulation via LLM Reasoning
Zhi Jing, Siyuan Yang, Jicong Ao +3
For robotic manipulation, existing robotics datasets and simulation benchmarks predominantly cater to robot-arm platforms. However, for humanoid robots equipped with dual arms and…
From Local Cues to Global Percepts: Emergent Gestalt Organization in Self-Supervised Vision Models
Tianqin Li, Ziqi Wen, Leiran Song +3
Human vision organizes local cues into coherent global forms using Gestalt principles like closure, proximity, and figure-ground assignment -- functions reliant on global spatial s…