6 papers
MedCutMix: A Data-Centric Approach to Improve Radiology Vision-Language Pre-training with Disease Awareness
Sinuo Wang, Yutong Xie, Yuyuan Liu +1
Vision-Language Pre-training (VLP) is drawing increasing interest for its ability to minimize manual annotation requirements while enhancing semantic understanding in downstream ta…
NavBench: Probing Multimodal Large Language Models for Embodied Navigation
Yanyuan Qiao, Haodong Hong, Wenqi Lyu +5
Multimodal Large Language Models (MLLMs) have demonstrated strong generalization in vision-language tasks, yet their ability to understand and act within embodied environments rema…
MMCLIP: Cross-modal Attention Masked Modelling for Medical Language-Image Pre-Training
Biao Wu, Yutong Xie, Zeyu Zhang +4
Vision-and-language pretraining (VLP) in the medical field utilizes contrastive learning on image-text pairs to achieve effective transfer across tasks. Yet, current VLP approaches…
A Comprehensive Analysis of Mamba for 3D Volumetric Medical Image Segmentation
Chaohan Wang, Yutong Xie, Qi Chen +2
Mamba, with its selective State Space Models (SSMs), offers a more computationally efficient solution than Transformers for long-range dependency modeling. However, there is still…
PathoHR: Breast Cancer Survival Prediction on High-Resolution Pathological Images
Yang Luo, Shiru Wang, Jun Liu +7
Breast cancer survival prediction in computational pathology presents a remarkable challenge due to tumor heterogeneity. For instance, different regions of the same tumor in the pa…
UD-Mamba: A pixel-level uncertainty-driven Mamba model for medical image segmentation
Weiren Zhao, Feng Wang, Yanran Wang +3
Recent advancements have highlighted the Mamba framework, a state-space model known for its efficiency in capturing long-range dependencies with linear computational complexity. Wh…