1 citations · 1 across the 1 of their papers we have counts for
6 papers
Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning
Kyle Domico, Jean-Charles Noirot Ferrand, Ryan Sheatsley +3
Attacks on machine learning models have been extensively studied through stateless optimization. In this paper, we demonstrate how a reinforcement learning (RL) agent can learn a n…
Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs
Jean-Charles Noirot Ferrand, Yohan Beugin, Eric Pauley +2
Alignment in large language models (LLMs) is used to enforce guidelines such as safety. Yet, alignment fails in the face of jailbreak attacks that modify inputs to induce unsafe ou…
On the Robustness Tradeoff in Fine-Tuning
Kunyang Li, Jean-Charles Noirot Ferrand, Ryan Sheatsley +4
Fine-tuning has become the standard practice for adapting pre-trained models to downstream tasks. However, the impact on model robustness is not well understood. In this work, we c…
Efficient Storage Integrity in Adversarial Settings
Quinn Burke, Ryan Sheatsley, Yohan Beugin +4
Storage integrity is essential to systems and applications that use untrusted storage (e.g., public clouds, end-user devices). However, known methods for achieving storage integrit…
Err on the Side of Texture: Texture Bias on Real Data
Blaine Hoak, Ryan Sheatsley, Patrick McDaniel
Bias significantly undermines both the accuracy and trustworthiness of machine learning models. To date, one of the strongest biases observed in image classification models is text…
On Scalable Integrity Checking for Secure Cloud Disks
Quinn Burke, Ryan Sheatsley, Rachel King +3
Merkle hash trees are the standard method to protect the integrity and freshness of stored data. However, hash trees introduce additional compute and I/O costs on the I/O critical…