38 citations · 51 across the 10 of their papers we have counts for
10 papers
What If We Recaption Billions of Web Images with LLaMA-3?
Xianhang Li, Haoqin Tu, Mude Hui +9
Web-crawled image-text pairs are inherently noisy. Prior studies demonstrate that semantically aligning and enriching textual descriptions of these pairs can significantly enhance…
Autoregressive Pretraining with Mamba in Vision
Sucheng Ren, Xianhang Li, Haoqin Tu +9
The vision community has started to build with the recently developed state space model, Mamba, as the new backbone for a range of tasks. This paper shows that Mamba's visual capab…
Medical Vision Generalist: Unifying Medical Imaging Tasks in Context
Sucheng Ren, Xiaoke Huang, Xianhang Li +5
This study presents Medical Vision Generalist (MVG), the first foundation model capable of handling various medical imaging tasks -- such as cross-modal synthesis, image segmentati…
3D-TransUNet for Brain Metastases Segmentation in the BraTS2023 Challenge
Siwei Yang, Xianhang Li, Jieru Mei +3
Segmenting brain tumors is complex due to their diverse appearances and scales. Brain metastases, the most common type of brain tumor, are a frequent complication of cancer. Theref…
SPFormer: Enhancing Vision Transformer with Superpixel Representation
Jieru Mei, Liang-Chieh Chen, Alan Yuille +1
In this work, we introduce SPFormer, a novel Vision Transformer enhanced by superpixel representation. Addressing the limitations of traditional Vision Transformers' fixed-size, no…
3D TransUNet: Advancing Medical Image Segmentation through Vision Transformers
Jieneng Chen, Jieru Mei, Xianhang Li +12
Medical image segmentation plays a crucial role in advancing healthcare systems for disease diagnosis and treatment planning. The u-shaped architecture, popularly known as U-Net, h…