1 paper
Robert Wolfe, Aylin Caliskan
We use a dataset of U.S. first names with labels based on predominant gender and racial group to examine the effect of training corpus frequency on tokenization, contextualization,…