2 papers
cs.AI2026
Recalling Too Well: Sycophancy Evaluation and Mitigation in Memory-Augmented Models
Shelly Bensal, Axel Magnuson, Aparna Balagopalan +1
Persistent memory systems promise to make LLMs more helpful by storing user beliefs over time. We show they also make models less correct by amplifying sycophancy, wherein models p…
cs.CL2025
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Shelly Bensal, Umar Jamil, Christopher Bryant +5
We explore a method for improving the performance of large language models through self-reflection and reinforcement learning. By incentivizing the model to generate better self-re…