collaborators

6 papers

cs.LG2026

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

Priyank Agrawal, Ankur Samanta, Shervin Ghasemlou +4

Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot gene…

cs.AI2026

Structure Enables Effective Self-Localization of Errors in LLMs

Ankur Samanta, Akshayaa Magesh, Ayush Jain +8

Self-correction in language models remains elusive. In this work, we explore whether language models can explicitly localize errors in incorrect reasoning, as a path toward buildin…

cs.AI2026

Credit Assignment with Resets in Language Model Reasoning

Ankur Samanta, Akshayaa Magesh, Ayush Jain +7

Contemporary reinforcement learning with verifiable reward methods post-train language models on multi-step reasoning by assigning a single outcome reward uniformly across all toke…

cs.LG2025

Aligning LLMs Toward Multi-Turn Conversational Outcomes Using Iterative PPO

Daniel R. Jiang, Jalaj Bhandari, Yukai Yang +2

Optimizing large language models (LLMs) for multi-turn conversational outcomes remains a significant challenge, especially in goal-oriented settings like AI marketing or sales agen…

cs.SE2025

A Note on Code Quality Score: LLMs for Maintainable Large Codebases

Sherman Wong, Jalaj Bhandari, Leo Zhou Fan Yang +6

Maintaining code quality in large-scale software systems presents significant challenges, particularly in settings where a large numbers of engineers work concurrently on a codebas…

cs.LG2025

Aligned Multi Objective Optimization

Yonathan Efroni, Ben Kretzu, Daniel Jiang +4

To date, the multi-objective optimization literature has mainly focused on conflicting objectives, studying the Pareto front, or requiring users to balance tradeoffs. Yet, in machi…