149 citations · 405 across the 11 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2023★ 7 cited
TouchStone: Evaluating Vision-Language Models by Language Models
Shuai Bai, Shusheng Yang, Jinze Bai +6
Large vision-language models (LVLMs) have recently witnessed rapid advancements, exhibiting a remarkable capacity for perceiving, understanding, and processing visual information b…
cs.CV2023★ 149 cited
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Jinze Bai, Shuai Bai, Shusheng Yang +6
In this work, we introduce the Qwen-VL series, a set of large-scale vision-language models (LVLMs) designed to perceive and understand both texts and images. Starting from the Qwen…
cs.CV2023★ 42 cited
ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
Peng Wang, Shijie Wang, Junyang Lin +5
In this work, we explore a scalable way for building a general representation model toward unlimited modalities. We release ONE-PEACE, a highly extensible model with 4B parameters…