Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training
Yangyi Chen, Hao Peng, Tong Zhang +1
In standard large vision-language models (LVLMs) pre-training, the model typically maximizes the joint probability of the caption conditioned on the image via next-token prediction…
cs.CV2024
SOLO: A Single Transformer for Scalable Vision-Language Modeling
Yangyi Chen, Xingyao Wang, Hao Peng +1
We present SOLO, a single transformer for Scalable visiOn-Language mOdeling. Current large vision-language models (LVLMs) such as LLaVA mostly employ heterogeneous architectures th…