1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Yuxuan Li, Hirokazu Shirado, Sauvik Das
While advances in fairness and alignment have helped mitigate overt biases exhibited by large language models (LLMs) when explicitly prompted, we hypothesize that these models may…