7 papers · 1 filter
Does Your ViT Still Need U-Net for Segmentation?
Xin Li, Wenhui Zhu, Xuanzhao Dong +6
Medical image segmentation is dominated by U-Net-style encoder-decoder architectures. Vision Transformers (ViTs) overcome the limited receptive field of convolutional networks thro…
Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning
Xuanzhao Dong, Wenhui Zhu, Peijie Qiu +11
Despite their popularity and success, Multimodal Large Language Models (MLLMs) often struggle to interpret images accurately, which limits their reasoning capability in complex sce…
Hierarchical Mesh Transformers with Topology-Guided Pretraining for Morphometric Analysis of Brain Structures
Yujian Xiong, Mohammad Farazi, Yanxi Chen +8
Representation learning on large-scale unstructured volumetric and surface meshes poses significant challenges in neuroimaging, especially when models must incorporate diverse vert…
RetinalGPT: A Retinal Clinical Preference Conversational Assistant Powered by Large Vision-Language Models
Wenhui Zhu, Xin Li, Xiwen Chen +8
Recently, Multimodal Large Language Models (MLLMs) have gained significant attention for their remarkable ability to process and analyze non-textual data, such as images, videos, a…
Plasma-CycleGAN: Plasma Biomarker-Guided MRI to PET Cross-modality Translation Using Conditional CycleGAN
Yanxi Chen, Yi Su, Celine Dumitrascu +6
Cross-modality translation between MRI and PET imaging is challenging due to the distinct mechanisms underlying these modalities. Blood-based biomarkers (BBBMs) are revolutionizing…
Many-MobileNet: Multi-Model Augmentation for Robust Retinal Disease Classification
Hao Wang, Wenhui Zhu, Xuanzhao Dong +9
In this work, we propose Many-MobileNet, an efficient model fusion strategy for retinal disease classification using lightweight CNN architecture. Our method addresses key challeng…