1 citations · 1 across the 23 of their papers we have counts for
22 papers · 1 filter
The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning
Aakriti Agrawal, Souradip Chakraborty, Armin Saghafian +6
Process Reward Models (PRMs) improve credit assignment for reasoning by providing step-level feedback. However, we identify a hidden bias in PRMs caused by severe imbalance in step…
RL with Learnable Textual Feedback: A Bilevel Approach
Utsav Singh, Sidhaarth Sredharan, Souradip Chakraborty +1
Reinforcement learning with verifiable rewards can improve LLM reasoning, but learning remains sample-inefficient when terminal rewards are sparse. This has motivated a growing lin…
TRAM: Test-Time Risk Adaptation with Mixture of Agents
Mohamad Fares El Hajj Chehade, Amrit Singh Bedi, Amy Zhang +1
Deployed reinforcement learning agents often face safety requirements that are specified only after training, such as new hazard maps, revised risk thresholds, or behavioral alignm…
Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach
Utsav Singh, Souradip Chakraborty, Wesley A. Suttle +6
Hierarchical reinforcement learning (HRL) enables agents to solve complex, long-horizon tasks by decomposing them into manageable sub-tasks. However, HRL methods face two fundament…
Improved Sample Complexity For Diffusion Model Training Without Empirical Risk Minimizer Access
Mudit Gaur, Prashant Trivedi, Sasidhar Kunapuli +2
Diffusion models have demonstrated state-of-the-art performance across vision, language, and scientific domains. Despite their empirical success, prior theoretical analyses of the…
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
Anas Barakat, Souradip Chakraborty, Khushbu Pahwa +1
Pass@k is a widely used performance metric for verifiable large language model tasks, including mathematical reasoning, code generation, and short-answer reasoning. It defines succ…