3 papers
cs.AI2026
Mitigating Misalignment Contagion by Steering with Implicit Traits
Maria Chang, Ronny Luss, Miao Liu +3
Language models (LMs) are increasingly used in high-stakes, multi-agent settings, where following instructions and maintaining value alignment are critical. Most alignment research…
cs.LG2025
Final-Model-Only Data Attribution with a Unifying View of Gradient-Based Methods
Dennis Wei, Inkit Padhi, Soumya Ghosh +3
Training data attribution (TDA) is concerned with understanding model behavior in terms of the training data. This paper draws attention to the common setting where one has access…
cs.CL2025
Cross-Examiner: Evaluating Consistency of Large Language Model-Generated Explanations
Danielle Villa, Maria Chang, Keerthiram Murugesan +2
Large Language Models (LLMs) are often asked to explain their outputs to enhance accuracy and transparency. However, evidence suggests that these explanations can misrepresent the…