1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Yangyi Chen, Hao Peng, Tong Zhang +1
In standard large vision-language models (LVLMs) pre-training, the model typically maximizes the joint probability of the caption conditioned on the image via next-token prediction…