Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs
Zhiyuan Yao, Zheren Fu, Zhixiao Zheng +3
Multimodal Large Language Models (MLLMs) are critically hampered by hallucination, generating content inconsistent with the provided image. In this paper, we identify an internal s…
cs.CV2025
Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models
Hou Xia, Zheren Fu, Fangcan Ling +4
Large video language models (LVLMs) have made notable progress in video understanding, spurring the development of corresponding evaluation benchmarks. However, existing benchmarks…