6 papers
DSO: Direct Steering Optimization for Bias Mitigation
Lucas Monteiro Paes, Nivedha Sivakumar, Yinong Oliver Wang +4
Generative models are often deployed to make decisions on behalf of users, such as vision-language models (VLMs) identifying which person in a room is a doctor to help visually imp…
Bias after Prompting: Persistent Discrimination in Large Language Models
Nivedha Sivakumar, Natalie Mackraz, Samira Khorshidi +4
A dangerous assumption that can be made from prior work on the bias transfer hypothesis (BTH) is that biases do not transfer from pre-trained large language models (LLMs) to adapte…
Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
Falaah Arif Khan, Nivedha Sivakumar, Yinong Oliver Wang +5
Large language models (LLMs) have achieved impressive performance, leading to their widespread adoption as decision-support tools in resource-constrained contexts like hiring and a…
Fairness Dynamics During Training
Krishna Patel, Nivedha Sivakumar, Barry-John Theobald +2
We investigate fairness dynamics during Large Language Model (LLM) training to enable the diagnoses of biases and mitigations through training interventions like early stopping; we…
Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
Yinong Oliver Wang, Nivedha Sivakumar, Falaah Arif Khan +6
The recent rapid adoption of large language models (LLMs) highlights the critical need for benchmarking their fairness. Conventional fairness metrics, which focus on discrete accur…
Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models
Natalie Mackraz, Nivedha Sivakumar, Samira Khorshidi +4
Large language models (LLMs) are increasingly being adapted to achieve task-specificity for deployment in real-world decision systems. Several previous works have investigated the…