Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
OVIBench: Benchmarking Online Video Question Answering under Interruption
Naiming Liu, Zhiheng Wu, Shuning Wang +3
Recent vision language models (VLMs) have achieved strong progress in video understanding. However, most existing video QA research and benchmarks still follow an offline, single-r…
cs.CV2026
Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding
Bowen Liu, Shuning Wang, Xinpeng Ding +3
Long-video understanding commonly compresses videos into a small set of frames or visual tokens for answer generation. Existing compact pipelines focus on retaining relevant visual…
cs.CV2026
See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding
Shuning Wang, Zhiheng Wu, YiNuo Lu +6
Recent advances in Video Large Language Models (Video-LLMs) have enabled performance on long-video understanding tasks. However, existing methods still face two key limitations: ev…