1 paper
Animesh Maheshwari, Divyansh Sahu, Nishit Verma
Vision-language models reliably name objects in a scene, but do they represent the 3D layout those objects inhabit? We introduce a 3,034-sample human-curated benchmark targeting th…