4 papers
PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation
Weixing Chen, Zhuoqian Feng, Yang Liu +6
Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The challenge stems from dense objec…
ST-VLA: Enabling 4D-Aware Spatiotemporal Understanding for General Robot Manipulation
You Wu, Zixuan Chen, Cunxu Ou +9
Robotic manipulation in open-world environments requires reasoning across semantics, geometry, and long-horizon action dynamics. Existing hierarchical Vision-Language-Action (VLA)…
Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation
Lingfeng Zhang, Yuecheng Liu, Zhanguang Zhang +16
Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have made them powerful tools in embodied navigation, enabling agents to leverage commonsense…
SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning
Yuecheng Liu, Dafeng Chi, Shiguang Wu +13
Spatial reasoning is an essential problem in embodied AI research. Efforts to enhance spatial reasoning abilities through supplementary spatial data and fine-tuning have proven lim…