activity
20162026
most citedBeyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

565 citations · 1.6k across the 82 of their papers we have counts for

collaborators
Showing 2020Show all

9 papers · 1 filter

cs.CL2020★ 51 cited

Learning from others' mistakes: Avoiding dataset biases without modeling them

Victor Sanh, Thomas Wolf, Yonatan Belinkov +1

State-of-the-art natural language processing (NLP) models often learn to model dataset biases and surface form correlations instead of features that target the intended underlying…

cs.CL2020

Analyzing Individual Neurons in Pre-trained Language Models

Nadir Durrani, Hassan Sajjad, Fahim Dalvi +1

While a lot of analysis has been carried to demonstrate linguistic knowledge captured by the representations learned within deep NLP models, very little attention has been paid tow…

eess.AS2020

Similarity Analysis of Self-Supervised Speech Representations

Yu-An Chung, Yonatan Belinkov, James Glass

Self-supervised speech representation learning has recently been a prosperous research topic. Many algorithms have been proposed for learning useful representations from large-scal…

cs.CL2020★ 7 cited

Probing Neural Dialog Models for Conversational Understanding

Abdelrhman Saleh, Tovly Deutsch, Stephen Casper +2

The predominant approach to open-domain dialog generation relies on end-to-end training of neural models on chat datasets. However, this approach provides little insight as to what…

cs.CL2020

The Sensitivity of Language Models and Humans to Winograd Schema Perturbations

Mostafa Abdou, Vinit Ravishankar, Maria Barrett +3

Large-scale pretrained language models are the major driving force behind recent improvements in performance on the Winograd Schema Challenge, a widely employed test of common sens…

cs.CL2020

Similarity Analysis of Contextual Word Representation Models

John M. Wu, Yonatan Belinkov, Hassan Sajjad +3

This paper investigates contextual word representation models from the lens of similarity analysis. Given a collection of trained models, we measure the similarity of their interna…