6 papers · 1 filter
MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory
Minghao Guo, Qingyue Jiao, Zeru Shi +14
Long-term agent memory is increasingly multimodal, yet existing evaluations rarely test whether agents preserve the visual evidence needed for later reasoning. In prior work, many…
Reinforcing Consistency in Video MLLMs with Structured Rewards
Yihao Quan, Zeru Shi, Jinman Zhao +1
Multimodal large language models (MLLMs) have achieved remarkable progress in video understanding. However, seemingly plausible outputs often suffer from poor visual and temporal g…
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
Liwei Che, Zhiyu Xue, Yihao Quan +7
Counting serves as a simple but powerful test of a Large Vision-Language Model's (LVLM's) reasoning; it forces the model to identify each individual object and then add them all up…
Improving Visual Reasoning with Iterative Evidence Refinement
Zeru Shi, Kai Mei, Yihao Quan +2
Vision language models (VLMs) are increasingly capable of reasoning over images, but robust visual reasoning often requires re-grounding intermediate steps in the underlying visual…
Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation
Xiaofeng Zhang, Fanshuo Zeng, Yihao Quan +2
Multimodal large language models have experienced rapid growth, and numerous different models have emerged. The interpretability of LVLMs remains an under-explored area. Especially…
Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
Xiaofeng Zhang, Yihao Quan, Chaochen Gu +6
The hallucination problem in multimodal large language models (MLLMs) remains a common issue. Although image tokens occupy a majority of the input sequence of MLLMs, there is limit…