7 papers
EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding
Yuejiao Su, Xinshen Zhang, Zhen Ye +3
Understanding human--environment interactions from egocentric vision is essential for assistive robotics and embodied intelligent agents, yet existing multimodal large language mod…
IGV-RRT: Prior-Real-Time Observation Fusion for Active Object Search in Changing Environments
Wei Zhang, Ping Gong, Yujie Wang +7
Object Goal Navigation (ObjectNav) in temporally changing indoor environments is challenging because object relocation can invalidate historical scene knowledge. To address this is…
HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding
Lei Yao, Yong Chen, Yuejiao Su +3
Humans commonly identify 3D object affordance through observed interactions in images or videos, and once formed, such knowledge can be generically generalized to novel objects. In…
Interaction-aware Representation Modeling with Co-occurrence Consistency for Egocentric Hand-Object Parsing
Yuejiao Su, Yi Wang, Lei Yao +2
A fine-grained understanding of egocentric human-environment interactions is crucial for developing next-generation embodied agents. One fundamental challenge in this area involves…
LaSSM: Efficient Semantic-Spatial Query Decoding via Local Aggregation and State Space Models for 3D Instance Segmentation
Lei Yao, Yi Wang, Yawen Cui +2
Query-based 3D scene instance segmentation from point clouds has attained notable performance. However, existing methods suffer from the query initialization dilemma due to the spa…
GVSynergy-Det: Synergistic Gaussian-Voxel Representations for Multi-View 3D Object Detection
Yi Zhang, Yi Wang, Lei Yao +1
Image-based 3D object detection aims to identify and localize objects in 3D space using only RGB images, eliminating the need for expensive depth sensors required by point cloud-ba…