3 papers
cs.CV2026
RoiMAM: Region-of-Interest Medical Attention Model for Efficient Vision-Language Understanding
Jiayan Yang, Zhuoyu Wu, Wenqi Fang
Vision-Language Models (VLMs) facilitate medical visual question answering (MedVQA) by jointly interpreting images and text. However, existing models typically depend on large arch…
eess.IV2026
EndoCaver: Handling Fog, Blur and Glare in Endoscopic Images via Joint Deblurring-Segmentation
Zhuoyu Wu, Wenhui Ou, Pei-Sze Tan +4
Endoscopic image analysis is vital for colorectal cancer screening, yet real-world conditions often suffer from lens fogging, motion blur, and specular highlights, which severely c…
eess.IV2025
RT-Focuser: A Real-Time Lightweight Model for Edge-side Image Deblurring
Zhuoyu Wu, Wenhui Ou, Qiawei Zheng +6
Motion blur caused by camera or object movement severely degrades image quality and poses challenges for real-time applications such as autonomous driving, UAV perception, and medi…