149 citations · 319 across the 3 of their papers we have counts for
1 paper · 1 filter
Jinze Bai, Shuai Bai, Shusheng Yang +6
In this work, we introduce the Qwen-VL series, a set of large-scale vision-language models (LVLMs) designed to perceive and understand both texts and images. Starting from the Qwen…