evidence routing 1high-resolution visual question answering 1memory efficiency 1multimodal large language models 1single-pass inference 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
Is It Time for the Renaissance of Salient Object Detection in the Era of MLLMs?
Wenzhuo Zhao, Xiuzhi Li, Zhongkuan Mao +6
The zero-shot capabilities of multimodal large language models (MLLMs) are pushing salient object detection (SOD) beyond task-specific supervision. To disentangle MLLMs beyond conv…
cs.CV2026
Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA
Zhongkuan Mao, Xianjie Liu, Tianyu Meng +9
The paper proposes a training‑free, single‑pass method that routes intermediate‑layer visual evidence to improve high‑resolution visual question answering without extra image proce…
cs.CV2026
Attend to Anything: Foundation Model for Unified Human Attention Modeling
Wenzhuo Zhao, Ronghao Xian, Keren Fu +1
Existing human attention (saliency) modeling methods persist as highly fragmented across modalities, scenes, and task formulations. Consequently, even with increasing model capacit…