activity
20242026
collaborators

9 papers

cs.CV2026

Does Your ViT Still Need U-Net for Segmentation?

Xin Li, Wenhui Zhu, Xuanzhao Dong +6

Medical image segmentation is dominated by U-Net-style encoder-decoder architectures. Vision Transformers (ViTs) overcome the limited receptive field of convolutional networks thro…

cs.CV2026

SGW-GAN: Sliced Gromov-Wasserstein Guided GANs for Retinal Fundus Image Enhancement

Yujian Xiong, Xuanzhao Dong, Wenhui Zhu +3

Retinal fundus photography is indispensable for ophthalmic screening and diagnosis, yet image quality is often degraded by noise, artifacts, and uneven illumination. Recent GAN- an…

cs.CV2025

nnMobileNet++: Towards Efficient Hybrid Networks for Retinal Image Analysis

Xin Li, Wenhui Zhu, Xuanzhao Dong +4

Retinal imaging is a critical, non-invasive modality for the early detection and monitoring of ocular and systemic diseases. Deep learning, particularly convolutional neural networ…

cs.CV2025

VAOT: Vessel-Aware Optimal Transport for Retinal Fundus Enhancement

Xuanzhao Dong, Wenhui Zhu, Yujian Xiong +8

Color fundus photography (CFP) is central to diagnosing and monitoring retinal disease, yet its acquisition variability (e.g., illumination changes) often degrades image quality, w…

cs.CV2025

A BERT-Style Self-Supervised Learning CNN for Disease Identification from Retinal Images

Xin Li, Wenhui Zhu, Peijie Qiu +3

In the field of medical imaging, the advent of deep learning, especially the application of convolutional neural networks (CNNs) has revolutionized the analysis and interpretation…

cs.CV2025

RetinalGPT: A Retinal Clinical Preference Conversational Assistant Powered by Large Vision-Language Models

Wenhui Zhu, Xin Li, Xiwen Chen +8

Recently, Multimodal Large Language Models (MLLMs) have gained significant attention for their remarkable ability to process and analyze non-textual data, such as images, videos, a…