7 papers · 1 filter
EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding
Yuejiao Su, Xinshen Zhang, Zhen Ye +3
Understanding human--environment interactions from egocentric vision is essential for assistive robotics and embodied intelligent agents, yet existing multimodal large language mod…
HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding
Lei Yao, Yong Chen, Yuejiao Su +3
Humans commonly identify 3D object affordance through observed interactions in images or videos, and once formed, such knowledge can be generically generalized to novel objects. In…
Interaction-aware Representation Modeling with Co-occurrence Consistency for Egocentric Hand-Object Parsing
Yuejiao Su, Yi Wang, Lei Yao +2
A fine-grained understanding of egocentric human-environment interactions is crucial for developing next-generation embodied agents. One fundamental challenge in this area involves…
LaSSM: Efficient Semantic-Spatial Query Decoding via Local Aggregation and State Space Models for 3D Instance Segmentation
Lei Yao, Yi Wang, Yawen Cui +2
Query-based 3D scene instance segmentation from point clouds has attained notable performance. However, existing methods suffer from the query initialization dilemma due to the spa…
GVSynergy-Det: Synergistic Gaussian-Voxel Representations for Multi-View 3D Object Detection
Yi Zhang, Yi Wang, Lei Yao +1
Image-based 3D object detection aims to identify and localize objects in 3D space using only RGB images, eliminating the need for expensive depth sensors required by point cloud-ba…
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
Lei Yao, Yi Wang, Yi Zhang +2
The significance of informative and robust point representations has been widely acknowledged for 3D scene understanding. Despite existing self-supervised pre-training counterparts…