4 papers
ATOM-Bench: A Real-World Benchmark for Atomic Skills and Compositional Generalization in Manipulation Policies
Zenan Wu, Bingqing Wei, Lu Liu +8
Generalist manipulation policies are increasingly presented as foundation models for robotic control, but their real-world generalization remains difficult to diagnose. A policy ma…
Feat2Go: Visual Feature-Grounded Value Estimation for Embodied Reinforcement Learning
Junyang Shu, Zhiwei Lin, Bingqing Wei +1
Reinforcement learning is a promising approach for improving the capabilities of vision-language-action (VLA) models while avoiding the heavy data requirements of imitation learnin…
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding
Zhongyu Xia, Yousen Tang, Bingqing Wei +1
Vision-Language-Action models have achieved remarkable progress in robotic manipulation, yet they suffer from a critical limitation: a lack of 3D scene understanding. This deficien…
ELITE: Experiential Learning and Intent-Aware Transfer for Self-improving Embodied Agents
Bingqing Wei, Zhongyu Xia, Dingai Liu +3
Vision-language models (VLMs) have shown remarkable general capabilities, yet embodied agents built on them fail at complex tasks, often skipping critical steps, proposing invalid…