7 papers
NutriBench-Kitchen: Benchmarking Embodied AI for Nutrition Management
Yulin Wei, Xiangchen Wang, Jianhui Pan +5
An embodied kitchen assistant must do more than recognize food in isolated frames. It must track ingredient states over time and integrate visual observations with recipe and nutri…
GVC-Seg: Training-Free 3D Instance Segmentation via Geometric Visual Correspondence
Liang Xu, Fangjing Wang, Jinyu Yang +1
Accurate 3D instance segmentation in point cloud data is critical for machine vision applications. Recent advancements leverage multiple pre-trained foundation models to generate 3…
CAPruner: Conceptual-Adjacent Scene Graph Pruner for Enhancing 3D Spatial Reasoning of Large Language Models
Shengli Zhou, Xiangchen Wang, Guanhua Chen +1
Large language models (LLMs) have recently been applied to 3D vision-language (3D-VL) tasks, which require spatial reasoning to identify target objects relative to anchors. Scene g…
Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models
Shengli Zhou, Minghang Zheng, Feng Zheng +1
Spatial reasoning focuses on locating target objects based on spatial relations in 3D scenes, which plays a crucial role in developing intelligent embodied agents. Due to the limit…
Learn 3D VQA Better with Active Selection and Reannotation
Shengli Zhou, Yang Liu, Feng Zheng
3D Visual Question Answering (3D VQA) is crucial for enabling models to perceive the physical world and perform spatial reasoning. In 3D VQA, the free-form nature of answers often…
: Toward Versatile Embodied Agents
Shengli Zhou, Xiangchen Wang, Jinrui Zhang +5
Embodied agents have demonstrated promising capabilities in interacting with physical environments. Yet, versatile embodied agents face three core bottlenecks: dynamic environmenta…