1 paper · 1 filter
Jagdish Tripathy, Marcus Buckmann
Instruction-tuned language models exhibit behavioural fairness in high-stakes decisions while retaining biased associations in their internal representations. However, whether thes…