Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
Rui Zhu, Xin Shen, Shuchen Wu +6
Spatial reasoning has emerged as a critical capability for Multimodal Large Language Models (MLLMs), drawing increasing attention and rapid advancement. However, existing benchmark…
cs.CV2025
R^3-VQA: "Read the Room" by Video Social Reasoning
Lixing Niu, Jiapeng Li, Xingping Yu +6
"Read the room" is a significant social reasoning capability in human daily life. Humans can infer others' mental states from subtle social cues. Previous social reasoning tasks an…
cs.CV2024
An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models
Fatemeh Shiri, Xiao-Yu Guo, Mona Golestan Far +3
Large Multimodal Models (LMMs) have achieved strong performance across a range of vision and language tasks. However, their spatial reasoning capabilities are under-investigated. I…