309 citations · 312 across the 9 of their papers we have counts for
3 papers · 2 filters
MixCE: Training Autoregressive Language Models by Mixing Forward and Reverse Cross-Entropies
Shiyue Zhang, Shijie Wu, Ozan Irsoy +4
Autoregressive language models are trained by minimizing the cross-entropy of the model distribution Q relative to the data distribution P -- that is, minimizing the forward cross-…
Distillation of encoder-decoder transformers for sequence labelling
Marco Farina, Duccio Pappadopulo, Anant Gupta +3
Driven by encouraging results on a wide range of tasks, the field of NLP is experiencing an accelerated race to develop bigger language models. This race for bigger models has also…
Weakly Supervised Headline Dependency Parsing
Adrian Benton, Tianze Shi, Ozan İrsoy +1
English news headlines form a register with unique syntactic properties that have been documented in linguistics literature since the 1930s. However, headlines have received surpri…