15 citations · 19 across the 7 of their papers we have counts for
10 papers
Inadvertent Context Leakage in Language Models
Jaiden Fairoze, Neal Mangaokar, Kamalika Chaudhuri +2
For AI agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere p…
Muse Spark Safety & Preparedness Report
Cristina Menghini, Peter Ney, Hamza Kwisaba +117
Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framewo…
CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
Niloofar Mireshghallah, Neal Mangaokar, Narine Kokhlikyan +4
Large Language Models (LLMs) increasingly use persistent memory from past interactions to enhance personalization and task performance. However, this memory introduces critical ris…
What Really is a Member? Discrediting Membership Inference via Poisoning
Neal Mangaokar, Ashish Hooda, Zhuohang Li +5
Membership inference tests aim to determine whether a particular data point was included in a language model's training set. However, recent works have shown that such tests often…
Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and Conditions in Shopping Websites at Scale
Elisa Tsai, Neal Mangaokar, Boyuan Zheng +2
Terms and conditions for online shopping websites often contain terms that can have significant financial consequences for customers. Despite their impact, there is currently no co…
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
Neal Mangaokar, Ashish Hooda, Jihye Choi +4
Large language models (LLMs) are typically aligned to be harmless to humans. Unfortunately, recent work has shown that such models are susceptible to automated jailbreak attacks th…