2 papers
cs.CV2026
Visual Token Compression Enhances Robustness of MLLMs
Shishen Gu, Jiequan Cui, Wenbo Hu +3
In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbrea…
cs.CV2025
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
Youze Wang, Zijun Chen, Ruoyu Chen +8
Recent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges…