activity
20152023
most citedBigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition

164 citations · 263 across the 19 of their papers we have counts for

collaborators

37 papers

cs.CV2023★ 2 cited

Less is More: Removing Text-regions Improves CLIP Training Efficiency and Robustness

Liangliang Cao, Bowen Zhang, Chen Chen +5

The CLIP (Contrastive Language-Image Pre-training) model and its variants are becoming the de facto backbone in many applications. However, training a CLIP model from hundreds of m…

cs.CV2023

STAIR: Learning Sparse Text and Image Representation in Grounded Tokens

Chen Chen, Bowen Zhang, Liangliang Cao +7

Image and text retrieval is one of the foundational tasks in the vision and language domain with multiple real-world applications. State-of-the-art approaches, e.g. CLIP, ALIGN, re…

cs.CV2022★ 1 cited

Exploiting Category Names for Few-Shot Classification with Vision-Language Models

Taihong Xiao, Zirui Wang, Liangliang Cao +3

Vision-language foundation models pretrained on large-scale data provide a powerful tool for many visual understanding tasks. Notably, many vision-language models build two encoder…

eess.AS2021★ 10 cited

Input Length Matters: Improving RNN-T and MWER Training for Long-form Telephony Speech Recognition

Zhiyun Lu, Yanwei Pan, Thibault Doutre +5

End-to-end models have achieved state-of-the-art results on several automatic speech recognition tasks. However, they perform poorly when evaluated on long-form data, e.g., minutes…

eess.AS2021★ 1 cited

Improving Confidence Estimation on Out-of-Domain Data for End-to-End Speech Recognition

Qiujia Li, Yu Zhang, David Qiu +3

As end-to-end automatic speech recognition (ASR) models reach promising performance, various downstream tasks rely on good confidence estimators for these systems. Recent research…

eess.AS2021★ 164 cited

BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition

Yu Zhang, Daniel S. Park, Wei Han +23

We summarize the results of a host of efforts using giant automatic speech recognition (ASR) models pre-trained using large, diverse unlabeled datasets containing approximately a m…