collaborators

5 papers

cs.CV2026

QuoVLA: Quotient Space for Vision-Language-Action Models

Xuan Wang, Yinan Wu, Haoran Duan +1

Vision-Language-Action (VLA) models commonly adapt pretrained Vision-Language Models (VLMs) to robot control by mapping visual observations and language instructions to continuous…

cs.AI2026

ConceptMoE: Concept-Guided Multimodal Mixture of Experts for Interpretable Computational Pathology

Xuan Wang, Zhongling Xu, Gopi Kannedhara +13

Healthcare models are transitioning from unimodal prediction toward multimodal reasoning over heterogeneous diagnostic inputs. In computational pathology, for complex tumor subtype…

cs.LG2026

Beyond the Aggregation Dilemma: Prior-Retaining Decoupled Learning for Multimodal Graphs

Hao Yan, Xuanru Wang, Jun Yin +3

Multimodal Attributed Graph Learning (MAGL) integrates intrinsic node attributes with structural topology via graph aggregation. However, as pretrained encoders evolve into Large F…

cs.CV2025

Enhancing Target-unspecific Tasks through a Features Matrix

Fangming Cui, Yonggang Zhang, Xuan Wang +2

Recent developments in prompt learning of large Vision-Language Models (VLMs) have significantly improved performance in target-specific tasks. However, these prompting methods oft…

cs.CV2025

Generalizable Prompt Learning of CLIP: A Brief Overview

Fangming Cui, Yonggang Zhang, Xuan Wang +2

Existing vision-language models (VLMs) such as CLIP have showcased an impressive capability to generalize well across various downstream tasks. These models leverage the synergy be…