49 citations · 84 across the 3 of their papers we have counts for
3 papers
cs.LG2020★ 49 cited
Weight Poisoning Attacks on Pre-trained Models
Keita Kurita, Paul Michel, Graham Neubig
Recently, NLP has seen a surge in the usage of large pre-trained models. Users download weights of models pre-trained on large datasets, then fine-tune the weights on a task of the…
cs.CL2019★ 16 cited
Towards Robust Toxic Content Classification
Keita Kurita, Anna Belova, Antonios Anastasopoulos
Toxic content detection aims to identify content that can offend or harm its recipients. Automated classifiers of toxic content need to be robust against adversaries who deliberate…
cs.CL2019★ 19 cited
Measuring Bias in Contextualized Word Representations
Keita Kurita, Nidhi Vyas, Ayush Pareek +2
Contextual word embeddings such as BERT have achieved state of the art performance in numerous NLP tasks. Since they are optimized to capture the statistical properties of training…