343 citations · 430 across the 14 of their papers we have counts for
5 papers · 1 filter
Optimizing Data Usage via Differentiable Rewards
Xinyi Wang, Hieu Pham, Paul Michel +3
To acquire a new skill, humans learn better and faster if a tutor, based on their current knowledge level, informs them of how much attention they should pay to particular content…
Findings of the First Shared Task on Machine Translation Robustness
Xian Li, Paul Michel, Antonios Anastasopoulos +7
We share the findings of the first shared task on improving robustness of Machine Translation (MT). The task provides a testbed representing challenges facing MT models deployed in…
Are Sixteen Heads Really Better than One?
Paul Michel, Omer Levy, Graham Neubig
Attention is a powerful and ubiquitous mechanism for allowing neural models to focus on particular salient pieces of information by taking their weighted average when making predic…
On Evaluation of Adversarial Perturbations for Sequence-to-Sequence Models
Paul Michel, Xian Li, Graham Neubig +1
Adversarial examples --- perturbations to the input of a model that elicit large changes in the output --- have been shown to be an effective way of assessing the robustness of seq…
compare-mt: A Tool for Holistic Comparison of Language Generation Systems
Graham Neubig, Zi-Yi Dou, Junjie Hu +4
In this paper, we describe compare-mt, a tool for holistic analysis and comparison of the results of systems for language generation tasks such as machine translation. The main goa…