activity
20242026
collaborators

11 papers

cs.CV2026

SGW-GAN: Sliced Gromov-Wasserstein Guided GANs for Retinal Fundus Image Enhancement

Yujian Xiong, Xuanzhao Dong, Wenhui Zhu +3

Retinal fundus photography is indispensable for ophthalmic screening and diagnosis, yet image quality is often degraded by noise, artifacts, and uneven illumination. Recent GAN- an…

cs.HC2026

EZBlender: Efficient 3D Editing with Plan-and-ReAct Agent

Hao Wang, Wenhui Zhu, Shao Tang +8

As a cornerstone of the modern digital economy, 3D modeling and rendering demand substantial resources and manual effort when scene editing is performed in the traditional manner.…

cs.SD2026

AHA: Aligning Large Audio-Language Models for Reasoning Hallucinations via Counterfactual Hard Negatives

Yanxi Chen, Wenhui Zhu, Xiwen Chen +9

Although Large Audio-Language Models (LALMs) deliver state-of-the-art (SOTA) performance, they frequently suffer from hallucinations, e.g. generating text not grounded in the audio…

cs.CV2025

nnMobileNet++: Towards Efficient Hybrid Networks for Retinal Image Analysis

Xin Li, Wenhui Zhu, Xuanzhao Dong +4

Retinal imaging is a critical, non-invasive modality for the early detection and monitoring of ocular and systemic diseases. Deep learning, particularly convolutional neural networ…

cs.CV2025

VAOT: Vessel-Aware Optimal Transport for Retinal Fundus Enhancement

Xuanzhao Dong, Wenhui Zhu, Yujian Xiong +8

Color fundus photography (CFP) is central to diagnosing and monitoring retinal disease, yet its acquisition variability (e.g., illumination changes) often degrades image quality, w…

cs.CL2025

Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models

Wenhui Zhu, Xuanzhao Dong, Xin Li +6

Recently, reinforcement learning (RL)-based tuning has shifted the trajectory of Multimodal Large Language Models (MLLMs), particularly following the introduction of Group Relative…