4 papers
Loss Landscape Poisoning: Targeted Extraction of Unseen Training Data from LLMs
Md Abdullah Al Mamun, Ngoc Phu Doan, Pedram Zaree +2
Large Language Models are increasingly trained on proprietary or sensitive data, from private healthcare and financial records to user conversations containing secrets. Ensuring th…
AttenMIA: LLM Membership Inference Attack through Attention Signals
Pedram Zaree, Md Abdullah Al Mamun, Yue Dong +2
Large Language Models (LLMs) are increasingly deployed to enable or improve a multitude of real-world applications. Given the large size of their training data sets, their tendency…
Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
Md Abdullah Al Mamun, Ihsen Alouani, Nael Abu-Ghazaleh
Large Language Models (LLMs) are aligned to meet ethical standards and safety requirements by training them to refuse answering harmful or unsafe prompts. In this paper, we demonst…
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
Pedram Zaree, Md Abdullah Al Mamun, Quazi Mishkatul Alam +3
Recent research has shown that carefully crafted jailbreak inputs can induce large language models to produce harmful outputs, despite safety measures such as alignment. It is impo…