87 citations · 170 across the 25 of their papers we have counts for
8 papers · 1 filter
Optimizing Deeper Transformers on Small Datasets
Peng Xu, Dhruv Kumar, Wei Yang +6
It is a common belief that training deep transformers from scratch requires large datasets. Consequently, for small datasets, people usually use shallow and simple additional layer…
Deconstructing word embedding algorithms
Kian Kenyon-Dean, Edward Newell, Jackie Chi Kit Cheung
Word embeddings are reliable feature representations of words used to obtain high quality results for various NLP applications. Uncontextualized word embeddings are used in many NL…
An Analysis of Dataset Overlap on Winograd-Style Tasks
Ali Emami, Adam Trischler, Kaheer Suleman +1
The Winograd Schema Challenge (WSC) and variants inspired by it have become important benchmarks for common-sense reasoning (CSR). Model performance on the WSC has quickly progress…
Learning Efficient Task-Specific Meta-Embeddings with Word Prisms
Jingyi He, KC Tsiolis, Kian Kenyon-Dean +1
Word embeddings are trained to predict word cooccurrence statistics, which leads them to possess different lexical properties (syntactic, semantic, etc.) depending on the notion of…
TeMP: Temporal Message Passing for Temporal Knowledge Graph Completion
Jiapeng Wu, Meng Cao, Jackie Chi Kit Cheung +1
Inferring missing facts in temporal knowledge graphs (TKGs) is a fundamental and challenging task. Previous works have approached this problem by augmenting methods for static know…
Multi-Fact Correction in Abstractive Text Summarization
Yue Dong, Shuohang Wang, Zhe Gan +3
Pre-trained neural abstractive summarization systems have dominated extractive strategies on news summarization performance, at least in terms of ROUGE. However, system-generated a…