5 papers
VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning
Mingkang Dong, Hongyi Cai, Jie Li +4
The effectiveness of multimodal instruction tuning depends not only on dataset scale, but critically on whether training samples genuinely require visual reasoning. However, existi…
Can Unified Generation and Understanding Models Maintain Semantic Equivalence Across Different Output Modalities?
Hongbo Jiang, Jie Li, Yunhang Shen +4
Unified Multimodal Large Language Models (U-MLLMs) integrate understanding and generation within a single architecture. However, existing evaluations typically assess these capabil…
AutoDebias: Automated Framework for Debiasing Text-to-Image Models
Hongyi Cai, Mohammad Mahdinur Rahman, Mingkang Dong +7
Text-to-Image (T2I) models generate high-quality images but are vulnerable to malicious backdoor attacks that inject harmful biases (e.g., trigger-activated gender or racial stereo…
Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
Yuan Xiong, Ziqi Miao, Lijun Li +3
While Multimodal Large Language Models (MLLMs) show remarkable capabilities, their safety alignments are susceptible to jailbreak attacks. Existing attack methods typically focus o…
STaR-Attack: A Spatio-Temporal and Narrative Reasoning Attack Framework for Unified Multimodal Understanding and Generation Models
Shaoxiong Guo, Tianyi Du, Lijun Li +3
Unified Multimodal understanding and generation Models (UMMs) have demonstrated remarkable capabilities in both understanding and generation tasks. However, we identify a vulnerabi…