collaborators

6 papers

cs.CR2026

Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning

Kyle Domico, Jean-Charles Noirot Ferrand, Ryan Sheatsley +3

Attacks on machine learning models have been extensively studied through stateless optimization. In this paper, we demonstrate how a reinforcement learning (RL) agent can learn a n…

cs.CR2026

Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs

Jean-Charles Noirot Ferrand, Yohan Beugin, Eric Pauley +2

Alignment in large language models (LLMs) is used to enforce guidelines such as safety. Yet, alignment fails in the face of jailbreak attacks that modify inputs to induce unsafe ou…

cs.LG2025

On the Robustness Tradeoff in Fine-Tuning

Kunyang Li, Jean-Charles Noirot Ferrand, Ryan Sheatsley +4

Fine-tuning has become the standard practice for adapting pre-trained models to downstream tasks. However, the impact on model robustness is not well understood. In this work, we c…

cs.CV2025

Alignment and Adversarial Robustness: Are More Human-Like Models More Secure?

Blaine Hoak, Kunyang Li, Patrick McDaniel

A small but growing body of work has shown that machine learning models which better align with human vision have also exhibited higher robustness to adversarial examples, raising…

cs.CV2025

On Synthetic Texture Datasets: Challenges, Creation, and Curation

Blaine Hoak, Patrick McDaniel

The influence of textures on machine learning models has been an ongoing investigation, specifically in texture bias/learning, interpretability, and robustness. However, due to the…

cs.CV2025

Err on the Side of Texture: Texture Bias on Real Data

Blaine Hoak, Ryan Sheatsley, Patrick McDaniel

Bias significantly undermines both the accuracy and trustworthiness of machine learning models. To date, one of the strongest biases observed in image classification models is text…