2 papers
cs.CL2026
Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
Jingyi Sun, Pepa Atanasova, Sagnik Ray Choudhury +2
Context utilisation, the ability of Language Models (LMs) to incorporate relevant information from the provided context when generating responses, remains largely opaque to users,…
cs.CL2026
BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
Sekh Mainul Islam, Nadav Borenstein, Siddhesh Milind Pawar +3
Understanding biases and stereotypes encoded in the weights of Large Language Models (LLMs) is crucial for developing effective mitigation strategies. However, biased behaviour is…