6 papers
The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning
Will Hawkins, Kaivalya Rawal, Jonathan Rystrøm +8
Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task. However, prior work has shown that this increase in capability…
Evaluating the Ability of Explanations to Disambiguate Models in a Rashomon Set
Kaivalya Rawal, Eoin Delaney, Zihao Fu +2
Explainable artificial intelligence (XAI) is concerned with producing explanations indicating the inner workings of models. For a Rashomon set of similarly performing models, expla…
FairImagen: Post-Processing for Bias Mitigation in Text-to-Image Models
Zihao Fu, Ryan Brown, Shun Shao +3
Text-to-image diffusion models, such as Stable Diffusion, have demonstrated remarkable capabilities in generating high-quality and diverse images from natural language prompts. How…
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
Harry Mayne, Ryan Othniel Kearns, Yushi Yang +4
To collaborate effectively with humans, language models must be able to explain their decisions in natural language. We study a specific type of self-explanation: self-generated co…
Evaluating Model Explanations without Ground Truth
Kaivalya Rawal, Zihao Fu, Eoin Delaney +1
There can be many competing and contradictory explanations for a single model prediction, making it difficult to select which one to use. Current explanation evaluation frameworks…
Resource-constrained Fairness
Sofie Goethals, Eoin Delaney, Brent Mittelstadt +1
Access to resources strongly constrains the decisions we make. While we might wish to offer every student a scholarship, or schedule every patient for follow-up meetings with a spe…