5 papers
Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention
Himanshu Singh, Ziwei Xu, A. V. Subramanyam +1
Large Language Models (LLMs) are powerful text generators, yet they can produce toxic or harmful content even when given seemingly harmless prompts. This presents a serious safety…
LLMs Can Unlearn Refusal with Only 1,000 Benign Samples
Yangyang Guo, Ziwei Xu, Si Liu +2
This study reveals a previously unexplored vulnerability in the safety alignment of Large Language Models (LLMs). Existing aligned LLMs predominantly respond to unsafe queries with…
Reasoning LLMs are Wandering Solution Explorers
Jiahao Lu, Ziwei Xu, Mohan Kankanhalli
Large Language Models (LLMs) have demonstrated impressive reasoning abilities through test-time computation (TTC) techniques such as chain-of-thought prompting and tree-based reaso…
Bullying the Machine: How Personas Increase LLM Vulnerability
Ziwei Xu, Udit Sanghi, Mohan Kankanhalli
Large Language Models (LLMs) are increasingly deployed in interactions where they are prompted to adopt personas. This paper investigates whether such persona conditioning affects…
Privacy Risks and Preservation Methods in Explainable Artificial Intelligence: A Scoping Review
Sonal Allana, Mohan Kankanhalli, Rozita Dara
Explainable Artificial Intelligence (XAI) has emerged as a pillar of Trustworthy AI and aims to bring transparency in complex models that are opaque by nature. Despite the benefits…