108 citations · 137 across the 29 of their papers we have counts for
6 papers · 1 filter
Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents
Or Zion Eliav, Eyal Lenga, Shir Bernstien +1
Traditional pentesting uses reconnaissance at each step to uncover unseen weaknesses, build stronger attacks, and advance the objective; we argue that AI agents require the same tr…
PRISM: Recovering Instruction Sets from Language Model Activations
Gilad Gressel, Rahul Pankajakshan, Julia Diament +3
As LLMs are deployed as agents, reliable monitoring requires knowing not only what they output, but which instructions are steering their behavior. This is difficult when models in…
LeakBoost: Perceptual-Loss-Based Membership Inference Attack
Amit Kravchik Taub, Fred M. Grabovski, Guy Amit +1
Membership inference attacks (MIAs) aim to determine whether a sample was part of a model's training set, posing serious privacy risks for modern machine-learning systems. Existing…
GAVEL: Towards Rule-Based Safety Through Activation Monitoring
Shir Rozenfeld, Rahul Pankajakshan, Itay Zloczower +3
Large language models (LLMs) are increasingly paired with activation-based monitoring to detect and prevent harmful behaviors that may not be apparent at the surface-text level. Ho…
Discussion Paper: The Threat of Real Time Deepfakes
Guy Frankovits, Yisroel Mirsky
Generative deep learning models are able to create realistic audio and video. This technology has been used to impersonate the faces and voices of individuals. These ``deepfakes''…
The Threat of Offensive AI to Organizations
Yisroel Mirsky, Ambra Demontis, Jaidip Kotak +7
AI has provided us with the ability to automate tasks, extract information from vast amounts of data, and synthesize media that is nearly indistinguishable from the real thing. How…