1 paper
Jiacheng Ye, Shansan Gong, Jiahui Gao +6
While autoregressive Large Vision-Language Models (VLMs) have achieved remarkable success, their sequential generation often limits their efficacy in complex visual planning and dy…