activity
20242026
collaborators

6 papers

cs.LG2026

DSO: Direct Steering Optimization for Bias Mitigation

Lucas Monteiro Paes, Nivedha Sivakumar, Yinong Oliver Wang +4

Generative models are often deployed to make decisions on behalf of users, such as vision-language models (VLMs) identifying which person in a room is a doctor to help visually imp…

cs.CL2025

Bias after Prompting: Persistent Discrimination in Large Language Models

Nivedha Sivakumar, Natalie Mackraz, Samira Khorshidi +4

A dangerous assumption that can be made from prior work on the bias transfer hypothesis (BTH) is that biases do not transfer from pre-trained large language models (LLMs) to adapte…

cs.CL2025

Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution

Falaah Arif Khan, Nivedha Sivakumar, Yinong Oliver Wang +5

Large language models (LLMs) have achieved impressive performance, leading to their widespread adoption as decision-support tools in resource-constrained contexts like hiring and a…

cs.CL2025

Fairness Dynamics During Training

Krishna Patel, Nivedha Sivakumar, Barry-John Theobald +2

We investigate fairness dynamics during Large Language Model (LLM) training to enable the diagnoses of biases and mitigations through training interventions like early stopping; we…

cs.CL2025

Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs

Yinong Oliver Wang, Nivedha Sivakumar, Falaah Arif Khan +6

The recent rapid adoption of large language models (LLMs) highlights the critical need for benchmarking their fairness. Conventional fairness metrics, which focus on discrete accur…

cs.CL2024

Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models

Natalie Mackraz, Nivedha Sivakumar, Samira Khorshidi +4

Large language models (LLMs) are increasingly being adapted to achieve task-specificity for deployment in real-world decision systems. Several previous works have investigated the…