3 papers
cs.CV2026
OVIBench: Benchmarking Online Video Question Answering under Interruption
Naiming Liu, Zhiheng Wu, Shuning Wang +3
Recent vision language models (VLMs) have achieved strong progress in video understanding. However, most existing video QA research and benchmarks still follow an offline, single-r…
cs.CV2026
See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection
Zhiheng Wu, Tong Wang, Shuning Wang +2
Recent advances in Vision-Language Models (VLMs) have benefited from Reinforcement Learning (RL) for enhanced reasoning. However, existing methods still face critical limitations,…
cs.CV2026
MIRAGE: A Micro-Interaction Relational Architecture for Grounded Exploration in Multi-Figure Artworks
Jui-Cheng Chiu, Yu-Chao Wang, Shengyang Luo +4
Appreciating multi-figure paintings requires understanding how characters relate through subtle cues like gaze alignment, gesture, and spatial arrangement. We present MIRAGE, an ev…