Showing 2026Show all
2 papers · 1 filter
cs.CL2026
How Value Induction Reshapes LLM Behaviour
Arnav Arora, Natalie Schluter, Katherine Metcalf +1
Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and empathy, and values, such as h…
cs.CL2026
BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
Sekh Mainul Islam, Nadav Borenstein, Siddhesh Milind Pawar +3
Understanding biases and stereotypes encoded in the weights of Large Language Models (LLMs) is crucial for developing effective mitigation strategies. However, biased behaviour is…