activity
20242026
collaborators

31 papers

cs.CV2026

Does Your ViT Still Need U-Net for Segmentation?

Xin Li, Wenhui Zhu, Xuanzhao Dong +6

Medical image segmentation is dominated by U-Net-style encoder-decoder architectures. Vision Transformers (ViTs) overcome the limited receptive field of convolutional networks thro…

cs.AI2026

S-SPPO: Semantic-Calibrated Self-Play Preference Optimization

Xiwen Chen, Wenhui Zhu, Jingjing Wang +13

Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preference Optimization (DPO). However, the standard Bradley-Terry instantiation of DPO…

cs.CL2026

DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning

Xiwen Chen, Wenhui Zhu, Peijie Qiu +7

Post-training LLMs with Reinforcement Learning, specifically Group Relative Policy Optimization (GRPO), has emerged as a paradigm for enhancing mathematical reasoning. However, sta…

cs.CV2026

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning

Xuanzhao Dong, Wenhui Zhu, Peijie Qiu +11

Despite their popularity and success, Multimodal Large Language Models (MLLMs) often struggle to interpret images accurately, which limits their reasoning capability in complex sce…

cs.CV2026

OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models

Xuanzhao Dong, Wenhui Zhu, Xiwen Chen +13

The advancement of general medical Multimodal Large Language Models (MLLMs) has shown great potential for building conversational assistants to support clinical diagnosis. However,…

cs.CV2026

Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models

Xiao Liu, Jiaxiang Liu, Boci Peng +6

Vision Language Models adapt well to downstream tasks but are highly vulnerable to adversarial perturbations that disrupt cross-modal semantic alignment. Existing defenses are larg…