activity
20202026
most citedStateful Defenses for Machine Learning Models Are Not Yet Secure Against Black-box Attacks

15 citations · 19 across the 7 of their papers we have counts for

collaborators

10 papers

cs.LG2026

Inadvertent Context Leakage in Language Models

Jaiden Fairoze, Neal Mangaokar, Kamalika Chaudhuri +2

For AI agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere p…

cs.CY2026

Muse Spark Safety & Preparedness Report

Cristina Menghini, Peter Ney, Hamza Kwisaba +117

Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framewo…

cs.CR2025

CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs

Niloofar Mireshghallah, Neal Mangaokar, Narine Kokhlikyan +4

Large Language Models (LLMs) increasingly use persistent memory from past interactions to enhance personalization and task performance. However, this memory introduces critical ris…

cs.LG2025

What Really is a Member? Discrediting Membership Inference via Poisoning

Neal Mangaokar, Ashish Hooda, Zhuohang Li +5

Membership inference tests aim to determine whether a particular data point was included in a language model's training set. However, recent works have shown that such tests often…

cs.CR2025★ 3 cited

Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and Conditions in Shopping Websites at Scale

Elisa Tsai, Neal Mangaokar, Boyuan Zheng +2

Terms and conditions for online shopping websites often contain terms that can have significant financial consequences for customers. Despite their impact, there is currently no co…

cs.CR2024

PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails

Neal Mangaokar, Ashish Hooda, Jihye Choi +4

Large language models (LLMs) are typically aligned to be harmless to humans. Unfortunately, recent work has shown that such models are susceptible to automated jailbreak attacks th…