1 citations · 1 across the 1 of their papers we have counts for
Showing cs.CRShow all
2 papers · 1 filter
cs.CR2026★ 1 cited
Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning
Kyle Domico, Jean-Charles Noirot Ferrand, Ryan Sheatsley +3
Attacks on machine learning models have been extensively studied through stateless optimization. In this paper, we demonstrate how a reinforcement learning (RL) agent can learn a n…
cs.CR2026
Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs
Jean-Charles Noirot Ferrand, Yohan Beugin, Eric Pauley +2
Alignment in large language models (LLMs) is used to enforce guidelines such as safety. Yet, alignment fails in the face of jailbreak attacks that modify inputs to induce unsafe ou…