4 papers · 1 filter
PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
Ziwen Li, Xin Wang, Hanlue Zhang +8
The Vision-Language-Action (VLA) models have demonstrated remarkable performance on embodied tasks and shown promising potential for real-world applications. However, current VLAs…
MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation
Jiaxin Huang, Runnan Chen, Ziwen Li +5
Reasoning segmentation aims to segment target objects in complex scenes based on human intent and spatial reasoning. While recent multimodal large language models (MLLMs) have demo…
SURPRISE3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes
Jiaxin Huang, Ziwen Li, Hanlve Zhang +6
The integration of language and 3D perception is critical for embodied AI and robotic systems to perceive, understand, and interact with the physical world. Spatial reasoning, a ke…
Open-Vocabulary Segmentation with Unpaired Mask-Text Supervision
Zhaoqing Wang, Xiaobo Xia, Ziye Chen +4
Current state-of-the-art open-vocabulary segmentation methods typically rely on image-mask-text triplet annotations for supervision. However, acquiring such detailed annotations is…