8 papers
Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection
Keertana Chidambaram, Andrew Ilyas, Vasilis Syrgkanis
Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model "actor" is inspected by a "monitor" (often another language model) for signs of…
Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback
Keertana Chidambaram, Sanath Kumar Krishnamurthy, Qiuling Xu +2
In recommendation systems, users interact with only a small fraction of a vast item catalog, producing feedback that is both sparse and noisy. This challenges post-training generat…
Pigeonholing: how bad prompts hurt models, causing collapse and mistakes
Hyunji Nam, Keertana Chidambaram, Dorottya Demszky +1
While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode collapse, a phenomenon we call…
Robust Post-Training for Generative Recommenders: Why Exponential Reward-Weighted SFT Outperforms RLHF
Keertana Chidambaram, Sanath Kumar Krishnamurthy, Qiuling Xu +2
Aligning generative recommender systems to user preferences via post-training is critical for closing the gap between next-item prediction and actual recommendation quality. Existi…
Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
Keertana Chidambaram, Karthik Vinary Seetharaman, Vasilis Syrgkanis
Reinforcement Learning from Human Feedback (RLHF) has become central to aligning large language models with human values, typically by first learning a reward model from preference…
Personalized Adaptation via In-Context Preference Learning
Allison Lau, Younwoo Choi, Vahid Balazadeh +3
Reinforcement Learning from Human Feedback (RLHF) is widely used to align Language Models (LMs) with human preferences. However, existing approaches often neglect individual user p…