collaborators

5 papers

cs.CV2026

Goal-oriented Navigation Instruction Generation with Tour Video Priors

Fangdi Li, Juncheng Liao, Changxu Cheng +4

Navigation Instruction Generation (NIG) aims to produce step-by-step natural language instructions for navigation guidance. Existing studies primarily treat NIG as an auxiliary tas…

cs.RO2026

Self-Imitated Diffusion Policy for Efficient and Robust Visual Navigation

Runhua Zhang, Junyi Hou, Changxu Cheng +3

Diffusion policies (DP) have demonstrated significant potential in visual navigation by capturing diverse multi-modal trajectory distributions. However, standard imitation learning…

cs.CV2025

ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation

Lingfeng Wang, Hualing Lin, Senda Chen +5

While humans effortlessly draw visual objects and shapes by adaptively allocating attention based on their complexity, existing multimodal large language models (MLLMs) remain cons…

cs.RO2025

DiffAD: A Unified Diffusion Modeling Approach for Autonomous Driving

Tao Wang, Cong Zhang, Xingguang Qu +3

End-to-end autonomous driving (E2E-AD) has rapidly emerged as a promising approach toward achieving full autonomy. However, existing E2E-AD systems typically adopt a traditional mu…

cs.CV2025

HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model

Tao Wang, Changxu Cheng, Lingfeng Wang +2

The remarkable performance of large multimodal models (LMMs) has attracted significant interest from the image segmentation community. To align with the next-token-prediction parad…