6 papers
Personalization, Personas, and Forecasting in Value Alignment
James Wedgwood, Pratiksha Thaker, Neil Kale +1
LLM behavior may be conditioned by human identity in several ways: they may be asked to adapt to users, role-play populations, or forecast how people would answer value-laden quest…
Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety
Neil Kale, Rebecca Portnoff, Pratiksha Thaker +5
Modern artificial intelligence (AI) systems present profound new risks to child safety. AI is increasingly being misused to create AI-generated child sexual abuse material, facilit…
Membership Inference Attacks for Unseen Classes
Pratiksha Thaker, Neil Kale, Zhiwei Steven Wu +1
A key tool in developing safe AI models is \emph{data auditing}, i.e., using statistical tools to determine whether harmful content may have been used in the training data of a bla…
BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
Shengyuan Hu, Neil Kale, Pratiksha Thaker +3
Machine unlearning has the potential to improve the safety of large language models (LLMs) by removing sensitive or harmful information post hoc. A key challenge in unlearning invo…
Position: LLM Unlearning Benchmarks are Weak Measures of Progress
Pratiksha Thaker, Shengyuan Hu, Neil Kale +3
Unlearning methods have the potential to improve the privacy and safety of large language models (LLMs) by removing sensitive or harmful information post hoc. The LLM unlearning re…
Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning
Shengyuan Hu, Yiwei Fu, Zhiwei Steven Wu +1
Machine unlearning is a promising approach to mitigate undesirable memorization of training data in ML models. However, in this work we show that existing approaches for unlearning…