5 papers
TRACE: Temporal Retrieval with Anchored and Convergent Evidence for Long-Horizon Video Understanding
Pengyiang Liu, Junbo Niu, Xiaoyang Hu +4
A long-video answer is evidence-supported only when the frames decoded from the video cover every event the answer depends on. Existing evaluations score final-answer correctness o…
OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs
Yifei Li, Pengyiang Liu, Yuhang Zang +4
Multimodal agents in robotics, AR, and autonomous driving must reason about places and layouts from continuous egocentric streams, often using evidence outside the current view. Ex…
SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance
Pengyiang Liu, Zhongyue Shi, Hongye Hao +7
Video understanding requires models to continuously track and update world state during playback. Although existing benchmarks have advanced video understanding evaluation across m…
Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM
Peng Liu, Xiaoming Ren, Fengkai Liu +5
Recent advancements in image-to-video (I2V) generation have shown promising performance in conventional scenarios. However, these methods still encounter significant challenges whe…
H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding
Qi Wu, Quanlong Zheng, Yanhao Zhang +8
With the rapid development of multimodal models, the demand for assessing video understanding capabilities has been steadily increasing. However, existing benchmarks for evaluating…