4 papers
Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance
Bohdan Turbal, Blossom Metevier, Max Springer +1
Adversarial attacks on large language models have limited practical impact despite extensive research. Optimization-based attacks such as Greedy Coordinate Gradient (GCG) (Zou et a…
The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety
Max Springer, Chung Peng Lee, Blossom Metevier +5
Fine-tuning aligned language models on benign tasks unpredictably degrades safety guardrails, even when training data contains no harmful content and developers have no adversarial…
ElliCE: Efficient and Provably Robust Algorithmic Recourse via the Rashomon Sets
Bohdan Turbal, Iryna Voitsitska, Lesia Semenova
Machine learning models now influence decisions that directly affect people's lives, making it important to understand not only their predictions, but also how individuals could ac…
On Adversarial Robustness of Language Models in Transfer Learning
Bohdan Turbal, Anastasiia Mazur, Jiaxu Zhao +1
We investigate the adversarial robustness of LLMs in transfer learning scenarios. Through comprehensive experiments on multiple datasets (MBIB Hate Speech, MBIB Political Bias, MBI…