5 citations · 8 across the 2 of their papers we have counts for
1 paper · 1 filter
Yangyi Chen, Xingyao Wang, Hao Peng +1
We present SOLO, a single transformer for Scalable visiOn-Language mOdeling. Current large vision-language models (LVLMs) such as LLaVA mostly employ heterogeneous architectures th…