activity
20242026
collaborators

5 papers

cs.RO2026

SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical Planning

Zebin Han, Xudong Wang, Baichen Liu +5

Sequential-Horizon Vision-and-Language Navigation (SH-VLN) presents a challenging scenario where agents should sequentially execute multi-task navigation guided by complex, long-ho…

cs.CV2025

GMT: Effective Global Framework for Multi-Camera Multi-Target Tracking

Yihao Zhen, Mingyue Xu, Qiang Wang +4

Multi-Camera Multi-Target (MCMT) tracking aims to locate and associate the same targets across multiple camera views. Existing methods typically adopt a two-stage framework, involv…

cs.CV2025

Complementary Information Guided Occupancy Prediction via Multi-Level Representation Fusion

Rongtao Xu, Jinzhou Lin, Jialei Zhou +6

Camera-based occupancy prediction is a mainstream approach for 3D perception in autonomous driving, aiming to infer complete 3D scene geometry and semantics from 2D images. Almost…

cs.CV2025

All in One: Visual-Description-Guided Unified Point Cloud Segmentation

Zongyan Han, Mohamed El Amine Boudjoghra, Jiahua Dong +2

Unified segmentation of 3D point clouds is crucial for scene understanding, but is hindered by its sparse structure, limited annotations, and the challenge of distinguishing fine-g…

cs.CV2024

SimC3D: A Simple Contrastive 3D Pretraining Framework Using RGB Images

Jiahua Dong, Tong Wu, Rui Qian +1

The 3D contrastive learning paradigm has demonstrated remarkable performance in downstream tasks through pretraining on point cloud data. Recent advances involve additional 2D imag…