4 papers
Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards
Xi Li, Shu Zhao, Xiaohan Zou +6
Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding and reasoning. However, this arc…
When Vision-Language Models Judge Without Seeing: Exposing Informativeness Bias
Xiaohan Zou, Roshan Sridhar, Mohammadtaher Safarzadeh +1
The reliability of VLM-as-a-Judge is critical for the automatic evaluation of vision-language models (VLMs). Despite recent progress, our analysis reveals that VLM-as-a-Judge often…
CARV: A Diagnostic Benchmark for Compositional Analogical Reasoning in Multimodal LLMs
Yongkang Du, Xiaohan Zou, Minhao Cheng +1
Analogical reasoning tests a fundamental aspect of human cognition: mapping the relation from one pair of objects to another. Existing evaluations of this ability in multimodal lar…
Understanding and Rectifying Safety Perception Distortion in VLMs
Xiaohan Zou, Jian Kang, George Kesidis +1
Recent studies reveal that vision-language models (VLMs) become more susceptible to harmful requests and jailbreak attacks after integrating the vision modality, exhibiting greater…