9 papers
Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety
Neil Kale, Rebecca Portnoff, Pratiksha Thaker +5
Modern artificial intelligence (AI) systems present profound new risks to child safety. AI is increasingly being misused to create AI-generated child sexual abuse material, facilit…
Alignment Defends LLMs from Property Inference Attacks
Pengrun Huang, Chhavi Yadav, Ruihan Wu +1
Large language models (LLMs) are increasingly fine-tuned on domain-specific datasets that may contain sensitive, dataset-level properties. Recent work has shown that such dataset-l…
Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks
Kevin Kuo, Chhavi Yadav, Virginia Smith
Recent defenses for safeguarding open-weight large language models (LLMs) are intended to prevent adversarial usage. Underlying these defenses is an assumption that new harmful beh…
Curriculum Learning for Safety Alignment
Sandeep Kumar, Virginia Smith, Chhavi Yadav
Direct Preference Optimisation (DPO) is widely used for safety alignment in large language models. However, prior work shows it is brittle and exhibits poor out-of-distribution (OO…
Automated Concept Discovery for LLM-as-a-Judge Preference Analysis
James Wedgwood, Chhavi Yadav, Virginia Smith
Large Language Models (LLMs) are increasingly used as scalable evaluators of model outputs, but their preference judgments exhibit systematic biases and can diverge from human eval…
Can We Infer Confidential Properties of Training Data from LLMs?
Pengrun Huang, Chhavi Yadav, Kamalika Chaudhuri +1
Large language models (LLMs) are increasingly fine-tuned on domain-specific datasets to support applications in fields such as healthcare, finance, and law. These fine-tuning datas…