activity
20172024
most citedEvaluating Layers of Representation in Neural Machine Translation on Part-of-Speech and Semantic Tagging Tasks

82 citations · 310 across the 24 of their papers we have counts for

collaborators

38 papers

cs.CL2021

Debiasing Methods in Natural Language Understanding Make Bias More Accessible

Michael Mendelson, Yonatan Belinkov

Model robustness to bias is often determined by the generalization on carefully designed out-of-distribution datasets. Recent debiasing methods in natural language understanding (N…

cs.CL20211 cited

A Generative Approach for Mitigating Structural Biases in Natural Language Inference

Dimion Asael, Zachary Ziegler, Yonatan Belinkov

Many natural language inference (NLI) datasets contain biases that allow models to perform well by only using a biased subset of the input, without considering the remainder featur…

cs.CL20211 cited

Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models

Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann +3

Targeted syntactic evaluations have demonstrated the ability of language models to perform subject-verb agreement given difficult contexts. To elucidate the mechanisms by which the…

cs.CL202139 cited

Variational Information Bottleneck for Effective Low-Resource Fine-Tuning

Rabeeh Karimi Mahabadi, Yonatan Belinkov, James Henderson

While large-scale pretrained language models have obtained impressive results when fine-tuned on a wide variety of tasks, they still often suffer from overfitting in low-resource s…

cs.CL2021

Probing Classifiers: Promises, Shortcomings, and Advances

Yonatan Belinkov

Probing classifiers have emerged as one of the prominent methodologies for interpreting and analyzing deep neural network models of natural language processing. The basic idea is s…

cs.CL202051 cited

Learning from others' mistakes: Avoiding dataset biases without modeling them

Victor Sanh, Thomas Wolf, Yonatan Belinkov +1

State-of-the-art natural language processing (NLP) models often learn to model dataset biases and surface form correlations instead of features that target the intended underlying…