1 paper · 1 filter
Alberto G. Rodriguez Salgado
How do multimodal models solve visual spatial tasks -- through genuine planning, or through brute-force search in token space? We introduce \textsc{MazeBench}, a benchmark of 110 p…