activity
20182024
most citedGlance-and-Gaze Vision Transformer

33 citations · 100 across the 10 of their papers we have counts for

collaborators

12 papers

cs.CV20241 cited

Evaluating Multiview Object Consistency in Humans and Image Models

Tyler Bonnen, Stephanie Fu, Yutong Bai +5

We introduce a benchmark to directly evaluate the alignment between human observers and vision models on a 3D shape inference task. We leverage an experimental design from the cogn…

cs.CV20226 cited

Delving into Masked Autoencoders for Multi-Label Thorax Disease Classification

Junfei Xiao, Yutong Bai, Alan Yuille +1

Vision Transformer (ViT) has become one of the most popular neural architectures due to its great scalability, computational efficiency, and compelling performance in many vision t…

cs.CV20226 cited

Making Your First Choice: To Address Cold Start Problem in Vision Active Learning

Liangyu Chen, Yutong Bai, Siyu Huang +4

Active learning promises to improve annotation efficiency by iteratively selecting the most important data to be annotated first. However, we uncover a striking contradiction to th…

cs.CV2022

Fast AdvProp

Jieru Mei, Yucheng Han, Yutong Bai +5

Adversarial Propagation (AdvProp) is an effective way to improve recognition models, leveraging adversarial examples. Nonetheless, AdvProp suffers from the extremely slow training…

cs.CV2022

Point-Level Region Contrast for Object Detection Pre-Training

Yutong Bai, Xinlei Chen, Alexander Kirillov +2

In this work we present point-level region contrast, a self-supervised pre-training approach for the task of object detection. This approach is motivated by the two key factors in…

cs.CV202133 cited

Glance-and-Gaze Vision Transformer

Qihang Yu, Yingda Xia, Yutong Bai +3

Recently, there emerges a series of vision Transformers, which show superior performance with a more compact model size than conventional convolutional neural networks, thanks to t…