11 papers
SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation
Changyuan Wang, Chubin Zhang, Zhenyu Wu +8
Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on robotic manipulation benchmarks. How…
RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation
Chuanrui Zhang, Zhengxian Wu, Guanxing Lu +2
Learned world models hold significant potential as neural simulators for robotic manipulation. However, prevalent 2D video-based models inherently lack the spatial and kinematic re…
RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation
Yuquan Xue, Guanxing Lu, Zhenyu Wu +4
Vision-Language-Action (VLA) models have shown strong manipulation capability when trained with large-scale imitation learning datasets. However, these datasets that predominantly…
BiDexGrasp: Coordinated Bimanual Dexterous Grasps across Object Geometries and Sizes
Mu Lin, Yi-Lin Wei, Jiaxuan Chen +7
Bimanual dexterous grasping is a fundamental and promising area in robotics, yet its progress is constrained by the lack of comprehensive datasets and powerful generation models. I…
VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
Wenkai Guo, Guanxing Lu, Haoyuan Deng +3
Vision-Language-Action models (VLAs) achieve strong performance in general robotic manipulation tasks by scaling imitation learning. However, existing VLAs are limited to predictin…
Human-in-the-loop Online Rejection Sampling for Robotic Manipulation
Guanxing Lu, Rui Zhao, Haitao Lin +2
Reinforcement learning (RL) is widely used to produce robust robotic manipulation policies, but fine-tuning vision-language-action (VLA) models with RL can be unstable due to inacc…