1 paper
Zheng Lu, Mingqi Gao, Qinlei Xie +9
Current benchmarks for embodied vision-language planning inadvertently favor linguistic next-token prediction over physically grounded next-state reasoning. This rewards models tha…