collaborators

6 papers

cs.CV2025

MedCutMix: A Data-Centric Approach to Improve Radiology Vision-Language Pre-training with Disease Awareness

Sinuo Wang, Yutong Xie, Yuyuan Liu +1

Vision-Language Pre-training (VLP) is drawing increasing interest for its ability to minimize manual annotation requirements while enhancing semantic understanding in downstream ta…

cs.CV2025

NavBench: Probing Multimodal Large Language Models for Embodied Navigation

Yanyuan Qiao, Haodong Hong, Wenqi Lyu +5

Multimodal Large Language Models (MLLMs) have demonstrated strong generalization in vision-language tasks, yet their ability to understand and act within embodied environments rema…

cs.CV2025

MMCLIP: Cross-modal Attention Masked Modelling for Medical Language-Image Pre-Training

Biao Wu, Yutong Xie, Zeyu Zhang +4

Vision-and-language pretraining (VLP) in the medical field utilizes contrastive learning on image-text pairs to achieve effective transfer across tasks. Yet, current VLP approaches…

cs.CV2025

A Comprehensive Analysis of Mamba for 3D Volumetric Medical Image Segmentation

Chaohan Wang, Yutong Xie, Qi Chen +2

Mamba, with its selective State Space Models (SSMs), offers a more computationally efficient solution than Transformers for long-range dependency modeling. However, there is still…

eess.IV2025

PathoHR: Breast Cancer Survival Prediction on High-Resolution Pathological Images

Yang Luo, Shiru Wang, Jun Liu +7

Breast cancer survival prediction in computational pathology presents a remarkable challenge due to tumor heterogeneity. For instance, different regions of the same tumor in the pa…

eess.IV2025

UD-Mamba: A pixel-level uncertainty-driven Mamba model for medical image segmentation

Weiren Zhao, Feng Wang, Yanran Wang +3

Recent advancements have highlighted the Mamba framework, a state-space model known for its efficiency in capturing long-range dependencies with linear computational complexity. Wh…