4 papers
Mitigating Positional Leakage in 3D Masked Autoencoders for Robust Representation Learning
Xu Yan, Huiqun Wang, Chen Wang +2
Masked autoencoding has emerged as a prominent paradigm for self-supervised learning on 3D point clouds, achieving competitive performance across downstream tasks. Unlike its 2D co…
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
Zhenghao Chen, Huiqun Wang, Di Huang
Multimodal large language models (MLLMs) are increasingly being applied to spatial cognition tasks, where they are expected to understand and interact with complex environments. Mo…
CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Models
Nan Zhou, Huiqun Wang, Yaoyan Zheng +1
Multimodal large language models (MLLMs) achieve remarkable progress in cross-modal perception and reasoning, yet a fundamental question remains unresolved: should the vision encod…
Implicit Modeling for Transferability Estimation of Vision Foundation Models
Yaoyan Zheng, Huiqun Wang, Nan Zhou +1
Transferability estimation identifies the best pre-trained models for downstream tasks without incurring the high computational cost of full fine-tuning. This capability facilitate…