6 papers
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information
Priyank Agrawal, Ankur Samanta, Shervin Ghasemlou +4
Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot gene…
Structure Enables Effective Self-Localization of Errors in LLMs
Ankur Samanta, Akshayaa Magesh, Ayush Jain +8
Self-correction in language models remains elusive. In this work, we explore whether language models can explicitly localize errors in incorrect reasoning, as a path toward buildin…
Credit Assignment with Resets in Language Model Reasoning
Ankur Samanta, Akshayaa Magesh, Ayush Jain +7
Contemporary reinforcement learning with verifiable reward methods post-train language models on multi-step reasoning by assigning a single outcome reward uniformly across all toke…
Aligning LLMs Toward Multi-Turn Conversational Outcomes Using Iterative PPO
Daniel R. Jiang, Jalaj Bhandari, Yukai Yang +2
Optimizing large language models (LLMs) for multi-turn conversational outcomes remains a significant challenge, especially in goal-oriented settings like AI marketing or sales agen…
A Note on Code Quality Score: LLMs for Maintainable Large Codebases
Sherman Wong, Jalaj Bhandari, Leo Zhou Fan Yang +6
Maintaining code quality in large-scale software systems presents significant challenges, particularly in settings where a large numbers of engineers work concurrently on a codebas…
Aligned Multi Objective Optimization
Yonathan Efroni, Ben Kretzu, Daniel Jiang +4
To date, the multi-objective optimization literature has mainly focused on conflicting objectives, studying the Pareto front, or requiring users to balance tradeoffs. Yet, in machi…