5 papers · 1 filter
Mitigating Misalignment Contagion by Steering with Implicit Traits
Maria Chang, Ronny Luss, Miao Liu +3
Language models (LMs) are increasingly used in high-stakes, multi-agent settings, where following instructions and maintaining value alignment are critical. Most alignment research…
Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models
Huzaifa Arif, Keerthiram Murugesan, Ching-Yun Ko +3
We propose patching for large language models (LLMs) like software versions, a lightweight and modular approach for addressing safety vulnerabilities. While vendors release improve…
Context Attribution with Multi-Armed Bandit Optimization
Deng Pan, Keerthiram Murugesan, Ting Hua +2
Understanding which parts of the retrieved context contribute to a large language model's generated answer is essential for building interpretable and trustworthy retrieval-augment…
The Unlearning Mirage: A Dynamic Framework for Evaluating LLM Unlearning
Raj Sanjay Shah, Jing Huang, Keerthiram Murugesan +2
Unlearning in Large Language Models (LLMs) aims to enhance safety, mitigate biases, and comply with legal mandates, such as the right to be forgotten. However, existing unlearning…
Language Models Coupled with Metacognition Can Outperform Reasoning Models
Vedant Khandelwal, Francesca Rossi, Keerthiram Murugesan +4
Large language models (LLMs) excel in speed and adaptability across various reasoning tasks, but they often struggle when strict logic or constraint enforcement is required. In con…