activity
20212023
most citedQwen Technical Report

110 citations · 127 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CL2023110 cited

Qwen Technical Report

Jinze Bai, Shuai Bai, Yunfei Chu +45

Large language models (LLMs) have revolutionized the field of artificial intelligence, enabling natural language processing tasks that were previously thought to be exclusive to hu…

cs.CV20231 cited

ViTMatte: Boosting Image Matting with Pretrained Plain Vision Transformers

Jingfeng Yao, Xinggang Wang, Shusheng Yang +1

Recently, plain vision Transformers (ViTs) have shown impressive performance on various computer vision tasks, thanks to their strong modeling capacity and large-scale pretraining.…

cs.CV20231 cited

RILS: Masked Visual Reconstruction in Language Semantic Space

Shusheng Yang, Yixiao Ge, Kun Yi +4

Both masked image modeling (MIM) and natural language supervision have facilitated the progress of transferable visual pre-training. In this work, we seek the synergy between two p…

cs.CV20224 cited

Unleashing Vanilla Vision Transformer with Masked Image Modeling for Object Detection

Yuxin Fang, Shusheng Yang, Shijie Wang +3

We present an approach to efficiently and effectively adapt a masked image modeling (MIM) pre-trained vanilla Vision Transformer (ViT) for object detection, which is based on our t…

cs.CV20222 cited

Temporally Efficient Vision Transformer for Video Instance Segmentation

Shusheng Yang, Xinggang Wang, Yu Li +5

Recently vision transformer has achieved tremendous success on image-level visual recognition tasks. To effectively and efficiently model the crucial temporal information within a…

cs.LG2022

Relational Surrogate Loss Learning

Tao Huang, Zekang Li, Hua Lu +6

Evaluation metrics in machine learning are often hardly taken as loss functions, as they could be non-differentiable and non-decomposable, e.g., average precision and F1 score. Thi…