1 paper
Zhao Liu, Tian Xie, Xueru Zhang
Current social bias benchmarks for Large Language Models (LLMs) primarily rely on predefined question formats like multiple-choice, limiting their ability to reflect the complexity…