1 paper
Houcheng Jiang, Jiajun Fu, Junfeng Fang +4
Multimodal large language models are increasingly expected to perform thinking with images, yet existing visual latent reasoning methods still rely on explicit textual chain-of-tho…