papers

Publications (52)

cs.CV2019

Class-Balanced Loss Based on Effective Number of Samples

Yin Cui, Menglin Jia, Tsung-Yi Lin +2

With the rapid increase of large-scale, real-world datasets, it becomes critical to address the problem of long-tailed data distribution (i.e., a few classes account for most of th…

cs.CV2021

Towards a Unified Foundation Model: Jointly Pre-Training Transformers on Unpaired Images and Text

Qing Li, Boqing Gong, Yin Cui +4

In this paper, we explore the possibility of building a unified foundation model that can be adapted to both vision-only and text-only tasks. Starting from BERT and ViT, we design…

cs.CV2020

Rethinking Pre-training and Self-training

Barret Zoph, Golnaz Ghiasi, Tsung-Yi Lin +4

Pre-training is a dominant paradigm in computer vision. For example, supervised ImageNet pre-training is commonly used to initialize the backbones of object detection and segmentat…

cs.CV2022

Bridging the Gap Between Object Detection and User Intent via Query-Modulation

Marco Fornoni, Chaochao Yan, Liangchen Luo +5

When interacting with objects through cameras, or pictures, users often have a specific intent. For example, they may want to perform a visual search. With most object detection mo…

cs.CV2021

Exploring Temporal Granularity in Self-Supervised Video Representation Learning

Rui Qian, Yeqing Li, Liangzhe Yuan +7

This work presents a self-supervised learning framework named TeG to explore Temporal Granularity in learning video representations. In TeG, we sample a long clip from a video and…

cs.CV2021

VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text

Hassan Akbari, Liangzhe Yuan, Rui Qian +4

We present a framework for learning multimodal representations from unlabeled data using convolution-free Transformer architectures. Specifically, our Video-Audio-Text Transformer…