11 citations · 11 across the 1 of their papers we have counts for
4 papers
An Interpretability Illusion for BERT
Tolga Bolukbasi, Adam Pearce, Ann Yuan +4
We describe an "interpretability illusion" that arises when analyzing the BERT model. Activations of individual neurons in the network may spuriously appear to encode a single, sim…
Debiasing Embeddings for Reduced Gender Bias in Text Classification
Flavien Prost, Nithum Thain, Tolga Bolukbasi
(Bolukbasi et al., 2016) demonstrated that pretrained word embeddings can inherit gender bias from the data they were trained on. We investigate how this bias affects downstream cl…
Robust Text Classifier on Test-Time Budgets
Md Rizwan Parvez, Tolga Bolukbasi, Kai-Wei Chang +1
We propose a generic and interpretable learning framework for building robust text classification model that achieves accuracy comparable to full models under test-time budget cons…
Quantifying and Reducing Stereotypes in Word Embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Zou +2
Machine learning algorithms are optimized to model statistical properties of the training data. If the input data reflects stereotypes and biases of the broader society, then the o…