1 paper
Pengyu Li, Zhitao Gao, Lingling Zhang +4
Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly an order-of-magnitude inference co…