activity
20212023
most citedQwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

149 citations · 181 across the 11 of their papers we have counts for

collaborators
Showing 2023Show all

5 papers · 1 filter

cs.CV2023★ 7 cited

TouchStone: Evaluating Vision-Language Models by Language Models

Shuai Bai, Shusheng Yang, Jinze Bai +6

Large vision-language models (LVLMs) have recently witnessed rapid advancements, exhibiting a remarkable capacity for perceiving, understanding, and processing visual information b…

cs.CV2023★ 149 cited

Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Jinze Bai, Shuai Bai, Shusheng Yang +6

In this work, we introduce the Qwen-VL series, a set of large-scale vision-language models (LVLMs) designed to perceive and understand both texts and images. Starting from the Qwen…

cs.CV2023★ 1 cited

ViTMatte: Boosting Image Matting with Pretrained Plain Vision Transformers

Jingfeng Yao, Xinggang Wang, Shusheng Yang +1

Recently, plain vision Transformers (ViTs) have shown impressive performance on various computer vision tasks, thanks to their strong modeling capacity and large-scale pretraining.…

cs.CV2023

MobileInst: Video Instance Segmentation on the Mobile

Renhong Zhang, Tianheng Cheng, Shusheng Yang +8

Video instance segmentation on mobile devices is an important yet very challenging edge AI problem. It mainly suffers from (1) heavy computation and memory costs for frame-by-frame…

cs.CV2023★ 1 cited

RILS: Masked Visual Reconstruction in Language Semantic Space

Shusheng Yang, Yixiao Ge, Kun Yi +4

Both masked image modeling (MIM) and natural language supervision have facilitated the progress of transferable visual pre-training. In this work, we seek the synergy between two p…