1 paper
Sarah Alnegheimish, Alicia Guo, Yi Sun
Evaluation of biases in language models is often limited to synthetically generated datasets. This dependence traces back to the need for a prompt-style dataset to trigger specific…