3 papers
cs.CV2026
Video-HOCA: A Diagnostic Benchmark for Physical Anomaly Reasoning in Video-LLMs
Chang Liu, Yunfan Ye, Qingyang Zhou +5
We introduce Video-HOCA, a diagnostic benchmark for physical anomaly reasoning in videos. Video-HOCA uses an Ontological-Causal taxonomy to distinguish violations of an entity's ow…
cs.CV2025
ALLVB: All-in-One Long Video Understanding Benchmark
Xichen Tan, Yuanjing Luo, Yunfan Ye +2
From image to video understanding, the capabilities of Multi-modal LLMs (MLLMs) are increasingly powerful. However, most existing video understanding benchmarks are relatively shor…
cs.CV2025
RAG-Adapter: A Plug-and-Play RAG-enhanced Framework for Long Video Understanding
Xichen Tan, Yunfan Ye, Yuanjing Luo +3
Multi-modal Large Language Models (MLLMs) capable of video understanding are advancing rapidly. To effectively assess their video comprehension capabilities, long video understandi…