6 papers
InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization
Haoxiang Ma, Junhao Cai, Xiaoxu Xu +26
Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and the physical dynamics learned through future prediction. In pract…
FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model
Xiaoxu Xu, Hao Li, Jinhui Ye +7
Predictive foresight is important to intelligent embodied agents. Since the motor execution of a robot is intrinsically constrained by its visual perception of environmental geomet…
DBGroup: Dual-Branch Point Grouping for Weakly Supervised 3D Semantic Instance Segmentation
Xuexun Liu, Xiaoxu Xu, Qiudan Zhang +2
Weakly supervised 3D instance segmentation is essential for 3D scene understanding, especially as the growing scale of data and high annotation costs associated with fully supervis…
3D Weakly Supervised Semantic Segmentation via Class-Aware and Geometry-Guided Pseudo-Label Refinement
Xiaoxu Xu, Xuexun Liu, Jinlong Li +5
3D weakly supervised semantic segmentation (3D WSSS) aims to achieve semantic segmentation by leveraging sparse or low-cost annotated data, significantly reducing reliance on dense…
Weakly-Supervised 3D Visual Grounding based on Visual Language Alignment
Xiaoxu Xu, Yitian Yuan, Qiudan Zhang +4
Learning to ground natural language queries to target objects or regions in 3D point clouds is quite essential for 3D scene understanding. Nevertheless, existing 3D visual groundin…
LESS: Label-Efficient and Single-Stage Referring 3D Segmentation
Xuexun Liu, Xiaoxu Xu, Jinlong Li +4
Referring 3D Segmentation is a visual-language task that segments all points of the specified object from a 3D point cloud described by a sentence of query. Previous works perform…