13 papers
Safety Alignment of LMs via Non-cooperative Games
Anselm Paulus, Ilia Kulikov, Brandon Amos +4
Ensuring the safety of language models (LMs) while maintaining their usefulness remains a critical challenge in AI alignment. Current approaches rely on sequential adversarial trai…
Muse Spark Safety & Preparedness Report
Cristina Menghini, Peter Ney, Hamza Kwisaba +117
Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framewo…
Privacy Blur: Quantifying Privacy and Utility for Image Data Release
Saeed Mahloujifar, Narine Kokhlikyan, Chuan Guo +1
Image data collected in the wild often contains private information such as faces and license plates, and responsible data release must ensure that this information stays hidden. A…
CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
Niloofar Mireshghallah, Neal Mangaokar, Narine Kokhlikyan +4
Large Language Models (LLMs) increasingly use persistent memory from past interactions to enhance personalization and task performance. However, this memory introduces critical ris…
Z0-Inf: Zeroth Order Approximation for Data Influence
Narine Kokhlikyan, Kamalika Chaudhuri, Saeed Mahloujifar
A critical aspect of analyzing and improving modern machine learning systems lies in understanding how individual training examples influence a model's predictive behavior. Estimat…
RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
Yuxin Wen, Arman Zharmagambetov, Ivan Evtimov +4
Prompt injection poses a serious threat to the reliability and safety of LLM agents. Recent defenses against prompt injection, such as Instruction Hierarchy and SecAlign, have show…