1 citations · 1 across the 1 of their papers we have counts for
2 papers
cs.CL2025
Policy Optimization Prefers The Path of Least Resistance
Debdeep Sanyal, Aakash Sen Sharma, Dhruv Kumar +2
Policy optimization (PO) algorithms are used to refine Large Language Models for complex, multi-step reasoning. Current state-of-the-art pipelines enforce a strict think-then-answe…
cs.CL2025★ 1 cited
Nine Ways to Break Copyright Law and Why Our LLM Won't: A Fair Use Aligned Generation Framework
Aakash Sen Sharma, Debdeep Sanyal, Priyansh Srivastava +4
Large language models (LLMs) commonly risk copyright infringement by reproducing protected content verbatim or with insufficient transformative modifications, posing significant et…