1 paper
Gio Huh, Dhruv Sheth, Rayhan Zirvi +1
While Vision-Language Models (VLMs) excel in many areas, they struggle with complex spatial reasoning, which requires problem decomposition and strategic tool use. Fine-tuning smal…