5 papers
QuoVLA: Quotient Space for Vision-Language-Action Models
Xuan Wang, Yinan Wu, Haoran Duan +1
Vision-Language-Action (VLA) models commonly adapt pretrained Vision-Language Models (VLMs) to robot control by mapping visual observations and language instructions to continuous…
ConceptMoE: Concept-Guided Multimodal Mixture of Experts for Interpretable Computational Pathology
Xuan Wang, Zhongling Xu, Gopi Kannedhara +13
Healthcare models are transitioning from unimodal prediction toward multimodal reasoning over heterogeneous diagnostic inputs. In computational pathology, for complex tumor subtype…
Beyond the Aggregation Dilemma: Prior-Retaining Decoupled Learning for Multimodal Graphs
Hao Yan, Xuanru Wang, Jun Yin +3
Multimodal Attributed Graph Learning (MAGL) integrates intrinsic node attributes with structural topology via graph aggregation. However, as pretrained encoders evolve into Large F…
Enhancing Target-unspecific Tasks through a Features Matrix
Fangming Cui, Yonggang Zhang, Xuan Wang +2
Recent developments in prompt learning of large Vision-Language Models (VLMs) have significantly improved performance in target-specific tasks. However, these prompting methods oft…
Generalizable Prompt Learning of CLIP: A Brief Overview
Fangming Cui, Yonggang Zhang, Xuan Wang +2
Existing vision-language models (VLMs) such as CLIP have showcased an impressive capability to generalize well across various downstream tasks. These models leverage the synergy be…