1 paper
Cheolhong Min, Jaeyun Jung, Daeun Lee +5
Vision-language models (VLMs) achieve strong performance on spatial reasoning benchmarks, yet it remains unclear whether this reflects structured 3D understanding or reliance on st…