9 papers
Privacy Blur: Quantifying Privacy and Utility for Image Data Release
Saeed Mahloujifar, Narine Kokhlikyan, Chuan Guo +1
Image data collected in the wild often contains private information such as faces and license plates, and responsible data release must ensure that this information stays hidden. A…
Hallucination reduction with CASAL: Contrastive Activation Steering For Amortized Learning
Wannan, Yang, Xinchi Qiu +6
Large Language Models (LLMs) exhibit impressive capabilities but often hallucinate, confidently providing incorrect answers instead of admitting ignorance. Prior work has shown tha…
CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
Niloofar Mireshghallah, Neal Mangaokar, Narine Kokhlikyan +4
Large Language Models (LLMs) increasingly use persistent memory from past interactions to enhance personalization and task performance. However, this memory introduces critical ris…
Z0-Inf: Zeroth Order Approximation for Data Influence
Narine Kokhlikyan, Kamalika Chaudhuri, Saeed Mahloujifar
A critical aspect of analyzing and improving modern machine learning systems lies in understanding how individual training examples influence a model's predictive behavior. Estimat…
RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
Yuxin Wen, Arman Zharmagambetov, Ivan Evtimov +4
Prompt injection poses a serious threat to the reliability and safety of LLM agents. Recent defenses against prompt injection, such as Instruction Hierarchy and SecAlign, have show…
How much do language models memorize?
John X. Morris, Chawin Sitawarin, Chuan Guo +5
We propose a new method for estimating how much a model knows about a datapoint and use it to measure the capacity of modern language models. Prior studies of language model memori…