4 papers
VTM-Nav: Harnessing Cross-Episode Experience for Object-Goal Navigation with Hierarchical Visual-Topological Memory
Xiaoran Xu, Yupeng Wu, Tianyu Xue +4
Training-free ObjectNav agents increasingly use vision-language models (VLMs), yet typically discard acquired scene knowledge after each request. We study cross-episode ObjectNav,…
DOME: Learning Transferable Domain Variables from Sparse Supervision for Test-Time Adaptation
Xiaoran Xu, Yifan Xu, Yupeng Wu +2
Test-time adaptation (TTA) aims to align a model to shifting test domains using only unlabeled streaming data. Most existing methods implicitly infer a single global domain distrib…
Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method
Xiaoran Xu, Xiaoshan Yang, Jiangang Yang +3
Open-Vocabulary Object Detection (OVOD) has achieved remarkable success in generalizing to novel categories. However, this success often rests on the implicit assumption of domain…
A Step Toward Federated Pretraining of Multimodal Large Language Models
Baochen Xiong, Yifan Xu, Xiaoshan Yang +3
The rapid evolution of Multimodal Large Language Models (MLLMs) is bottlenecked by the saturation of high-quality public data, while vast amounts of diverse multimodal data remain…