2 papers
cs.CV2026
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
Yuxiang Shen, Hailong Huang, Zhenkun Gao +6
Multimodal Large Language Models (MLLMs) are shifting towards "Thinking with Images" by actively exploring image details. While effective, large-scale training is computationally e…
cs.AI2026
TPRU: Advancing Temporal and Procedural Understanding in Large Multimodal Models
Zhenkun Gao, Xuhong Wang, Xin Tan +1
Multimodal Large Language Models (MLLMs), particularly smaller, deployable variants, exhibit a critical deficiency in understanding temporal and procedural visual data, a bottlenec…