3 papers
cs.CL2025
Bias after Prompting: Persistent Discrimination in Large Language Models
Nivedha Sivakumar, Natalie Mackraz, Samira Khorshidi +4
A dangerous assumption that can be made from prior work on the bias transfer hypothesis (BTH) is that biases do not transfer from pre-trained large language models (LLMs) to adapte…
cs.CL2025
Fairness Dynamics During Training
Krishna Patel, Nivedha Sivakumar, Barry-John Theobald +2
We investigate fairness dynamics during Large Language Model (LLM) training to enable the diagnoses of biases and mitigations through training interventions like early stopping; we…
cs.CL2024
Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models
Natalie Mackraz, Nivedha Sivakumar, Samira Khorshidi +4
Large language models (LLMs) are increasingly being adapted to achieve task-specificity for deployment in real-world decision systems. Several previous works have investigated the…