collaborators

13 papers

cs.AI2026

Safety Alignment of LMs via Non-cooperative Games

Anselm Paulus, Ilia Kulikov, Brandon Amos +4

Ensuring the safety of language models (LMs) while maintaining their usefulness remains a critical challenge in AI alignment. Current approaches rely on sequential adversarial trai…

cs.CY2026

Muse Spark Safety & Preparedness Report

Cristina Menghini, Peter Ney, Hamza Kwisaba +117

Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framewo…

cs.LG2025

Privacy Blur: Quantifying Privacy and Utility for Image Data Release

Saeed Mahloujifar, Narine Kokhlikyan, Chuan Guo +1

Image data collected in the wild often contains private information such as faces and license plates, and responsible data release must ensure that this information stays hidden. A…

cs.CR2025

CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs

Niloofar Mireshghallah, Neal Mangaokar, Narine Kokhlikyan +4

Large Language Models (LLMs) increasingly use persistent memory from past interactions to enhance personalization and task performance. However, this memory introduces critical ris…

cs.LG2025

Z0-Inf: Zeroth Order Approximation for Data Influence

Narine Kokhlikyan, Kamalika Chaudhuri, Saeed Mahloujifar

A critical aspect of analyzing and improving modern machine learning systems lies in understanding how individual training examples influence a model's predictive behavior. Estimat…

cs.CR2025

RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection

Yuxin Wen, Arman Zharmagambetov, Ivan Evtimov +4

Prompt injection poses a serious threat to the reliability and safety of LLM agents. Recent defenses against prompt injection, such as Instruction Hierarchy and SecAlign, have show…