activity
20172024
most citedImproving Alignment and Robustness with Circuit Breakers

7 citations · 24 across the 11 of their papers we have counts for

collaborators
Showing cs.CRShow all

5 papers · 1 filter

cs.CR20242 cited

Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents

Priyanshu Kumar, Elaine Lau, Saranya Vijayakumar +9

For safety reasons, large language models (LLMs) are trained to refuse harmful user instructions, such as assisting dangerous activities. We study an open question in this work: do…

cs.CR2024

LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses

Weiran Lin, Anna Gerchanovsky, Omer Akgul +3

Writing effective prompts for large language models (LLM) can be unintuitive and burdensome. In response, services that optimize or suggest prompts have emerged. While such service…

cs.CR2024

VeriSplit: Secure and Practical Offloading of Machine Learning Inferences across IoT Devices

Han Zhang, Zifan Wang, Mihir Dhamankar +2

Many Internet-of-Things (IoT) devices rely on cloud computation resources to perform machine learning inferences. This is expensive and may raise privacy concerns for users. Consum…

cs.CR20212 cited

The Design of the User Interfaces for Privacy Enhancements for Android

Jason I. Hong, Yuvraj Agarwal, Matt Fredrikson +22

We present the design and design rationale for the user interfaces for Privacy Enhancements for Android (PE for Android). These UIs are built around two core ideas, namely that dev…

cs.CR2018

Contextual and Granular Policy Enforcement in Database-backed Applications

Abhishek Bichhawat, Matt Fredrikson, Jean Yang +1

Database-backed applications rely on inlined policy checks to process users' private and confidential data in a policy-compliant manner as traditional database access control mecha…