4 papers
AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly
Zhi Jing, Jinbin Qiao, Ouyang Lu +5
Spatial reasoning is a fundamental capability for embodied intelligence, especially for fine-grained manipulation tasks such as robotic assembly. Recent methods based on vision-lan…
AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Afford Correspondence
Jiawei Zhang, Kaizhe Hu, Yingqian Huang +3
Despite the recent success of modern imitation learning methods in robot manipulation, their performance is often constrained by geometric variations due to limited data diversity.…
Beyond Action Residuals: Real-World Robot Policy Steering via Bottleneck Latent Reinforcement Learning
Dongjie Yu, Kun Lei, Zhennan Jiang +2
Pretrained imitation policies have become a strong foundation for robot manipulation, but they often require online improvement to overcome execution errors, limited dataset covera…
Dexora: Open-source VLA for High-DoF Bimanual Dexterity
Zongzheng Zhang, Jingrui Pang, Zhuo Yang +22
Vision-Language-Action (VLA) models have recently become a central direction in embodied AI, but current systems are restricted to either dual-gripper control or single-arm dextero…