1.2k citations · 1.3k across the 4 of their papers we have counts for
6 papers · 1 filter
Transcending Scaling Laws with 0.1% Extra Compute
Yi Tay, Jason Wei, Hyung Won Chung +13
Scaling language models improves performance but comes with significant computational costs. This paper proposes UL2R, a method that substantially improves existing language models…
Measuring and Reducing Gendered Correlations in Pre-trained Models
Kellie Webster, Xuezhi Wang, Ian Tenney +6
Pre-trained models have revolutionized natural language understanding. However, researchers have found they can encode artifacts undesired in many applications, such as professions…
Natural Language Processing with Small Feed-Forward Networks
Jan A. Botha, Emily Pitler, Ji Ma +5
We show that small and shallow feed-forward neural networks can achieve near state-of-the-art results on a range of unstructured and structured language processing tasks while bein…
SyntaxNet Models for the CoNLL 2017 Shared Task
Chris Alberti, Daniel Andor, Ivan Bogatyy +10
We describe a baseline dependency parsing system for the CoNLL2017 Shared Task. This system, which we call "ParseySaurus," uses the DRAGNN framework [Kong et al, 2017] to combine t…
Globally Normalized Transition-Based Neural Networks
Daniel Andor, Chris Alberti, David Weiss +5
We introduce a globally normalized transition-based neural network model that achieves state-of-the-art part-of-speech tagging, dependency parsing and sentence compression results.…
Structured Training for Neural Network Transition-Based Parsing
David Weiss, Chris Alberti, Michael Collins +1
We present structured perceptron training for neural network transition-based dependency parsing. We learn the neural network representation using a gold corpus augmented by a larg…