149 citations · 173 across the 9 of their papers we have counts for
8 papers · 1 filter
TouchStone: Evaluating Vision-Language Models by Language Models
Shuai Bai, Shusheng Yang, Jinze Bai +6
Large vision-language models (LVLMs) have recently witnessed rapid advancements, exhibiting a remarkable capacity for perceiving, understanding, and processing visual information b…
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Jinze Bai, Shuai Bai, Shusheng Yang +6
In this work, we introduce the Qwen-VL series, a set of large-scale vision-language models (LVLMs) designed to perceive and understand both texts and images. Starting from the Qwen…
ViTMatte: Boosting Image Matting with Pretrained Plain Vision Transformers
Jingfeng Yao, Xinggang Wang, Shusheng Yang +1
Recently, plain vision Transformers (ViTs) have shown impressive performance on various computer vision tasks, thanks to their strong modeling capacity and large-scale pretraining.…
RILS: Masked Visual Reconstruction in Language Semantic Space
Shusheng Yang, Yixiao Ge, Kun Yi +4
Both masked image modeling (MIM) and natural language supervision have facilitated the progress of transferable visual pre-training. In this work, we seek the synergy between two p…
Unleashing Vanilla Vision Transformer with Masked Image Modeling for Object Detection
Yuxin Fang, Shusheng Yang, Shijie Wang +3
We present an approach to efficiently and effectively adapt a masked image modeling (MIM) pre-trained vanilla Vision Transformer (ViT) for object detection, which is based on our t…
Temporally Efficient Vision Transformer for Video Instance Segmentation
Shusheng Yang, Xinggang Wang, Yu Li +5
Recently vision transformer has achieved tremendous success on image-level visual recognition tasks. To effectively and efficiently model the crucial temporal information within a…