2 papers
cs.CV2026
Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning
Chen Ling, Hanqian Li, Dongnan Liu +9
The visual reasoning ability of multimodal large language models (MLLMs) is crucial for downstream applications, particularly counter-commonsense reasoning, which requires models t…
cs.CV2026
Witness Evidence Portfolios: Single-Prefill Risk Detection for Closed Multimodal Answers
Fexiang Liu, Shiye Wang, Qiang Qiu +1
Reliable deployment of multimodal large language models (MLLMs) requires deciding whether a confident visual answer should be trusted, reviewed, or routed to a stronger system. Con…