9 papers
Dr. Jekyll and Mr. Hyde: Two Faces of LLMs
Matteo Gioele Collu, Tom Janssen-Groesbeek, Stefanos Koffas +2
Large Language Models (LLMs) are being integrated into applications such as chatbots or email assistants. To prevent improper responses, safety mechanisms, such as Reinforcement Le…
Backdoor Attacks on Decentralised Post-Training
OÄuzhan Ersoy, Nikolay Blagoev, Jona te Lintelo +3
Decentralised post-training of large language models utilises data and pipeline parallelism techniques to split the data and the model. Unfortunately, decentralised post-training c…
Removing the Trigger, Not the Backdoor: Alternative Triggers and Latent Backdoors
Gorka Abad, Ermes Franch, Stefanos Koffas +1
Current backdoor defenses assume that neutralizing a known trigger removes the backdoor. We show this trigger-centric view is incomplete: \emph{alternative triggers}, patterns perc…
Towards Backdoor Stealthiness in Model Parameter Space
Xiaoyun Xu, Zhuoran Liu, Stefanos Koffas +1
Recent research on backdoor stealthiness focuses mainly on indistinguishable triggers in input space and inseparable backdoor representations in feature space, aiming to circumvent…
SoK: The Last Line of Defense: On Backdoor Defense Evaluation
Gorka Abad, Marina KrÄek, Stefanos Koffas +7
Backdoor attacks pose a significant threat to deep learning models by implanting hidden vulnerabilities that can be activated by malicious inputs. While numerous defenses have been…
CatBack: Universal Backdoor Attacks on Tabular Data via Categorical Encoding
Behrad Tajalli, Stefanos Koffas, Stjepan Picek
Backdoor attacks in machine learning have drawn significant attention for their potential to compromise models stealthily, yet most research has focused on homogeneous data such as…