3 papers
cs.CV2026
TRACE: Temporal Retrieval with Anchored and Convergent Evidence for Long-Horizon Video Understanding
Pengyiang Liu, Junbo Niu, Xiaoyang Hu +4
A long-video answer is evidence-supported only when the frames decoded from the video cover every event the answer depends on. Existing evaluations score final-answer correctness o…
cs.CV2026
OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs
Yifei Li, Pengyiang Liu, Yuhang Zang +4
Multimodal agents in robotics, AR, and autonomous driving must reason about places and layouts from continuous egocentric streams, often using evidence outside the current view. Ex…
cs.CV2026
SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance
Pengyiang Liu, Zhongyue Shi, Hongye Hao +7
Video understanding requires models to continuously track and update world state during playback. Although existing benchmarks have advanced video understanding evaluation across m…