1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
XiuYu Zhang, Junfeng Fang, Zhenkai Liang
Latent visual reasoning (LVR) inserts supervised latent tokens between perception and answer generation in vision-language models (VLMs). The field uses alignment between these lat…