1 paper
Virginia K. Felkner, Ho-Chun Herbert Chang, Eugene Jang +1
We present WinoQueer: a benchmark specifically designed to measure whether large language models (LLMs) encode biases that are harmful to the LGBTQ+ community. The benchmark is com…