5 papers · 1 filter
AI Agents May Always Fall for Prompt Injections
Sahar Abdelnabi, Eugene Bagdasarian
Prompt injection is the most critical vulnerability in deployed AI agents. Despite recent progress, we show that the prevailing defense paradigm (data-instruction separation) both…
Contrastive Privacy: A Semantic Approach to Measuring Privacy of AI-based Sanitization
George Bissias, Eugene Bagdasarian, Brian Neil Levine
To sanitize specific concepts from imagery and text, privacy mechanisms with formal guarantees are often eschewed in practice in favor of more intuitive techniques. AI-based saniti…
Disappearing Ink: Obfuscation Breaks N-gram Code Watermarks in Theory and Practice
Gehao Zhang, Eugene Bagdasarian, Juan Zhai +1
Distinguishing AI-generated code from human-written code is becoming crucial for tasks such as authorship attribution, content tracking, and misuse detection. Based on this, N-gram…
Self-interpreting Adversarial Images
Tingwei Zhang, Collin Zhang, John X. Morris +2
We introduce a new type of indirect, cross-modal injection attacks against visual language models that enable creation of self-interpreting images. These images contain hidden "met…
Contextual Agent Security: A Policy for Every Purpose
Lillian Tsai, Eugene Bagdasarian
Judging an action's safety requires knowledge of the context in which the action takes place. To human agents who act in various contexts, this may seem obvious: performing an acti…