2 papers
cs.CV2026
Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA
Zhongkuan Mao, Xianjie Liu, Tianyu Meng +9
High-resolution visual question answering (HR-VQA) is often treated as a problem of insufficient evidence acquisition, where failing multimodal large language models must inspect i…
cs.CV2026
OmniReasoner: Thinking with Long Audio-Video via Native Tool Use
Yu Chen, Caorui Li, Ziyu Xiong +8
Long audio-video reasoning is difficult for omnimodal LLMs because the decisive evidence is often sparse, cross-modal, and too expensive to preserve with uniformly high-fidelity in…