activity
20242026
collaborators

11 papers

cs.CV2026

Visual Token Compression Enhances Robustness of MLLMs

Shishen Gu, Jiequan Cui, Wenbo Hu +3

In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbrea…

cs.AI2026

Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning

Zhen Zeng, Leijiang Gu, Zhangling Duan +4

Multimodal large language models (MLLMs) can inadvertently memorize privacy-sensitive information during training. While existing unlearning methods can remove such content, they o…

cs.CV2026

Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs

Ziqi Wang, Chang Che, Qi Wang +4

While continual visual instruction tuning (CVIT) has shown promise in adapting multimodal large language models (MLLMs), existing studies predominantly focus on models without safe…

cs.CV2026

Layer Consistency Matters: Elegant Latent Transition Discrepancy for Generalizable Synthetic Image Detection

Yawen Yang, Feng Li, Shuqi Kong +4

Recent rapid advancement of generative models has significantly improved the fidelity and accessibility of AI-generated synthetic images. While enabling various innovative applicat…

cs.IR2025

HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents

Anyang Tong, Xiang Niu, ZhiPing Liu +4

Existing multimodal Retrieval-Augmented Generation (RAG) methods for visually rich documents (VRD) are often biased towards retrieving salient knowledge(e.g., prominent text and vi…

cs.CV2025

Prompt to Restore, Restore to Prompt: Cyclic Prompting for Universal Adverse Weather Removal

Rongxin Liao, Feng Li, Yanyan Wei +4

Universal adverse weather removal (UAWR) seeks to address various weather degradations within a unified framework. Recent methods are inspired by prompt learning using pre-trained…