5 papers · 1 filter
The Realignment Problem: When Right becomes Wrong in LLMs
Aakash Sen Sharma, Debdeep Sanyal, Manodeep Ray +3
Post-training alignment of large language models (LLMs) relies on large-scale human annotations guided by policy specifications that change over time. Cultural shifts, value reinte…
Confidence is Not Competence
Debdeep Sanyal, Manya Pandey, Dhruv Kumar +2
Large language models (LLMs) often exhibit a puzzling disconnect between their asserted confidence and actual problem-solving competence. We offer a mechanistic account of this dec…
Policy Optimization Prefers The Path of Least Resistance
Debdeep Sanyal, Aakash Sen Sharma, Dhruv Kumar +2
Policy optimization (PO) algorithms are used to refine Large Language Models for complex, multi-step reasoning. Current state-of-the-art pipelines enforce a strict think-then-answe…
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
Debdeep Sanyal, Manodeep Ray, Murari Mandal
The release of open-weight large language models (LLMs) creates a tension between advancing accessible research and preventing misuse, such as malicious fine-tuning to elicit harmf…
Nine Ways to Break Copyright Law and Why Our LLM Won't: A Fair Use Aligned Generation Framework
Aakash Sen Sharma, Debdeep Sanyal, Priyansh Srivastava +4
Large language models (LLMs) commonly risk copyright infringement by reproducing protected content verbatim or with insufficient transformative modifications, posing significant et…