1 paper
Lixian Chen, Yanhui Chen, Mingxuan Huang +2
Vision--language models can face asymmetric visual and textual shifts at deployment. These shifts expose a multimodal failure mode in which an unreliable branch remains overconfide…