collaborators

16 papers

cs.CV2026

Localization-Infused Vision-Language Semantic Fusion for Text-Guided Medical Image Segmentation

Songyue Han, Mingye Zou, Shuchang Ye +2

Medical image segmentation is essential for modern computer-aided medicine. Recently, text-guided segmentation has shown promise by incorporating clinician-formulated textual repor…

cs.CV2026

EasyLens: A Training-Free Plug-and-Play Subtle-Lesion Representation Amplifier for Medical Vision-Language Models

Qiwei Zeng, Hao Wang, Jinghao Lin +6

Medical vision-language models (VLMs) have shown increasing potential for clinical image interpretation, including lesion detection and report generation. However, their practical…

cs.CV2026

Language-guided Medical Image Segmentation with Target-informed Multi-level Contrastive Alignments

Mingjian Li, Mingyuan Meng, Shuchang Ye +4

Medical image segmentation is a fundamental task in numerous medical engineering applications. Recently, language-guided segmentation has shown promise in medical scenarios where t…

cs.LG2026

From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation

Shuchang Ye, Jinqiang Yu, Zhujun Xiao +6

Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external APIs, including adaptation t…

cs.CE2026

ImProNCDE: Impulse-Corrected Neural Controlled Differential Equations with Prototype Learning for Longitudinal Prognosis Prediction

Hao Wang, Yupeng Xu, Jinghao Lin +5

Longitudinal ophthalmic imaging analysis is an essential step for prognosis prediction in ophthalmic diseases. However, AI-assisted prognosis models are challenged by follow-up seq…

cs.CL2026

Med-R2: Perception and Reflection-driven Complex Reasoning for Medical Report Generation

Hao Wang, Shuchang Ye, Jinghao Lin +2

Automated medical report generation (MRG) is increasingly used to reduce the burden of manual reporting and for decision support. Large vision-language models (LVLMs) hold great pr…