3 citations · 4 across the 8 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.CV2026
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
Zelin Xu, Yupu Zhang, Saugat Adhikari +6
Benchmarking spatial reasoning in multimodal large language models (MLLMs) has attracted growing interest in computer vision due to its importance for embodied AI and other agentic…
cs.LG2026
VNU-Bench: A Benchmarking Dataset for Multi-Source Multimodal News Video Understanding
Zibo Liu, Muyang Li, Zhe Jiang +1
News videos are carefully edited multimodal narratives that combine narration, visuals, and external quotations into coherent storylines. In recent years, there have been significa…