1 paper
Fatemeh Shiri, Xiao-Yu Guo, Mona Golestan Far +3
Large Multimodal Models (LMMs) have achieved strong performance across a range of vision and language tasks. However, their spatial reasoning capabilities are under-investigated. I…