activity
20242026
collaborators

5 papers

cs.LG2026

Theoretical Limits of Language Model Alignment

Lucas Monteiro Paes, Natalie Mackraz, Barry-John Theobald +1

Language model (LM) alignment improves model outputs to reflect human preferences while preserving the capabilities of the base model. The most common alignment approaches are (i)…

cs.CL2025

Bias after Prompting: Persistent Discrimination in Large Language Models

Nivedha Sivakumar, Natalie Mackraz, Samira Khorshidi +4

A dangerous assumption that can be made from prior work on the bias transfer hypothesis (BTH) is that biases do not transfer from pre-trained large language models (LLMs) to adapte…

cs.CL2025

Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs

Yinong Oliver Wang, Nivedha Sivakumar, Falaah Arif Khan +6

The recent rapid adoption of large language models (LLMs) highlights the critical need for benchmarking their fairness. Conventional fairness metrics, which focus on discrete accur…

cs.CL2025

Aligning LLMs by Predicting Preferences from User Writing Samples

Stéphane Aroca-Ouellette, Natalie Mackraz, Barry-John Theobald +1

Accommodating human preferences is essential for creating aligned LLM agents that deliver personalized and effective interactions. Recent work has shown the potential for LLMs acti…

cs.CL2024

Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models

Natalie Mackraz, Nivedha Sivakumar, Samira Khorshidi +4

Large language models (LLMs) are increasingly being adapted to achieve task-specificity for deployment in real-world decision systems. Several previous works have investigated the…