14 citations · 19 across the 4 of their papers we have counts for
4 papers
ExtremeBERT: A Toolkit for Accelerating Pretraining of Customized BERT
Rui Pan, Shizhe Diao, Jianlin Chen +1
In this paper, we present ExtremeBERT, a toolkit for accelerating and customizing BERT pretraining. Our goal is to provide an easy-to-use BERT pretraining toolkit for the research…
Normalizing Flow with Variational Latent Representation
Hanze Dong, Shizhe Diao, Weizhong Zhang +1
Normalizing flow (NF) has gained popularity over traditional maximum likelihood based methods due to its strong capability to model complex data distributions. However, the standar…
VLUE: A Multi-Task Benchmark for Evaluating Vision-Language Models
Wangchunshu Zhou, Yan Zeng, Shizhe Diao +1
Recent advances in vision-language pre-training (VLP) have demonstrated impressive performance in a range of vision-language (VL) tasks. However, there exist several challenges for…
ZEN: Pre-training Chinese Text Encoder Enhanced by N-gram Representations
Shizhe Diao, Jiaxin Bai, Yan Song +2
The pre-training of text encoders normally processes text as a sequence of tokens corresponding to small text units, such as word pieces in English and characters in Chinese. It om…