activity
20192026
most citedMedical Image Segmentation Using Squeeze-and-Expansion Transformers

24 citations · 90 across the 35 of their papers we have counts for

collaborators
Showing cs.CVShow all

14 papers · 1 filter

cs.CV2025

EVLF-FM: Explainable Vision Language Foundation Model for Medicine

Yang Bai, Haoran Cheng, Yang Zhou +40

Despite the promise of foundation models in medical AI, current systems remain limited - they are modality-specific and lack transparent reasoning processes, hindering clinical ado…

cs.CV2025

AdvMIM: Adversarial Masked Image Modeling for Semi-Supervised Medical Image Segmentation

Lei Zhu, Jun Zhou, Rick Siow Mong Goh +1

Vision Transformer has recently gained tremendous popularity in medical image segmentation task due to its superior capability in capturing long-range dependencies. However, transf…

cs.CV2025

Partially Supervised Unpaired Multi-Modal Learning for Label-Efficient Medical Image Segmentation

Lei Zhu, Yanyu Xu, Huazhu Fu +3

Unpaired Multi-Modal Learning (UMML) which leverages unpaired multi-modal data to boost model performance on each individual modality has attracted a lot of research interests in m…

cs.CV2024

Diffusion-Enhanced Test-time Adaptation with Text and Image Augmentation

Chun-Mei Feng, Yuanyang He, Jian Zou +6

Existing test-time prompt tuning (TPT) methods focus on single-modality data, primarily enhancing images and using confidence ratings to filter out inaccurate images. However, whil…

cs.CV2024

BenchX: A Unified Benchmark Framework for Medical Vision-Language Pretraining on Chest X-Rays

Yang Zhou, Tan Li Hui Faith, Yanyu Xu +4

Medical Vision-Language Pretraining (MedVLP) shows promise in learning generalizable and transferable visual representations from paired and unpaired medical images and reports. Me…

cs.CV2024★ 1 cited

From Generalist to Specialist: Adapting Vision Language Models via Task-Specific Visual Instruction Tuning

Yang Bai, Yang Zhou, Jun Zhou +3

Large vision language models (VLMs) combine large language models with vision encoders, demonstrating promise across various tasks. However, they often underperform in task-specifi…