25 citations · 89 across the 17 of their papers we have counts for
10 papers · 1 filter
ExtremeBERT: A Toolkit for Accelerating Pretraining of Customized BERT
Rui Pan, Shizhe Diao, Jianlin Chen +1
In this paper, we present ExtremeBERT, a toolkit for accelerating and customizing BERT pretraining. Our goal is to provide an easy-to-use BERT pretraining toolkit for the research…
MICO: A Multi-alternative Contrastive Learning Framework for Commonsense Knowledge Representation
Ying Su, Zihao Wang, Tianqing Fang +3
Commonsense reasoning tasks such as commonsense knowledge graph completion and commonsense question answering require powerful representation learning. In this paper, we propose to…
CONDA: a CONtextual Dual-Annotated dataset for in-game toxicity understanding and detection
Henry Weld, Guanghao Huang, Jean Lee +6
Traditional toxicity detection models have focused on the single utterance level without deeper understanding of context. We introduce CONDA, a new dataset for in-game toxic langua…
Neural Machine Translation with Adequacy-Oriented Learning
Xiang Kong, Zhaopeng Tu, Shuming Shi +2
Although Neural Machine Translation (NMT) models have advanced state-of-the-art performance in machine translation, they face problems like the inadequate translation. We attribute…
Multi-Head Attention with Disagreement Regularization
Jian Li, Zhaopeng Tu, Baosong Yang +2
Multi-head attention is appealing for the ability to jointly attend to information from different representation subspaces at different positions. In this work, we introduce a disa…
Modeling Localness for Self-Attention Networks
Baosong Yang, Zhaopeng Tu, Derek F. Wong +3
Self-attention networks have proven to be of profound value for its strength of capturing global dependencies. In this work, we propose to model localness for self-attention networ…