most citedSwin-UMamba: Mamba-based UNet with ImageNet-based pretraining

7 citations · 9 across the 4 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2024

Enhancing Representation in Medical Vision-Language Foundation Models via Multi-Scale Information Extraction Techniques

Weijian Huang, Cheng Li, Hong-Yu Zhou +6

The development of medical vision-language foundation models has attracted significant attention in the field of medicine and healthcare due to their promising prospect in various…

cs.CV20242 cited

MLIP: Medical Language-Image Pre-training with Masked Local Representation Learning

Jiarun Liu, Hong-Yu Zhou, Cheng Li +4

Existing contrastive language-image pre-training aims to learn a joint representation by matching abundant image-text pairs. However, the number of image-text pairs in medical data…

cs.CV2024

Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement

Weijian Huang, Cheng Li, Hao Yang +4

Recently, vision-language representation learning has made remarkable advancements in building up medical foundation models, holding immense potential for transforming the landscap…

cs.CV2024

A multi-modal vision-language model for generalizable annotation-free pathology localization

Hao Yang, Hong-Yu Zhou, Jiarun Liu +12

Existing deep learning models for defining pathology from clinical imaging data rely on expert annotations and lack generalization capabilities in open clinical environments. Here,…

cs.CV2024

Multimodal self-supervised learning for lesion localization

Hao Yang, Hong-Yu Zhou, Cheng Li +7

Multimodal deep learning utilizing imaging and diagnostic reports has made impressive progress in the field of medical imaging diagnostics, demonstrating a particularly strong capa…