activity
20242026
most citedAuction-Based Regulation for Artificial Intelligence

1 citations · 1 across the 23 of their papers we have counts for

collaborators
Showing cs.LGShow all

22 papers · 1 filter

cs.LG2026

The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning

Aakriti Agrawal, Souradip Chakraborty, Armin Saghafian +6

Process Reward Models (PRMs) improve credit assignment for reasoning by providing step-level feedback. However, we identify a hidden bias in PRMs caused by severe imbalance in step…

cs.LG2026

RL with Learnable Textual Feedback: A Bilevel Approach

Utsav Singh, Sidhaarth Sredharan, Souradip Chakraborty +1

Reinforcement learning with verifiable rewards can improve LLM reasoning, but learning remains sample-inefficient when terminal rewards are sparse. This has motivated a growing lin…

cs.LG2026

TRAM: Test-Time Risk Adaptation with Mixture of Agents

Mohamad Fares El Hajj Chehade, Amrit Singh Bedi, Amy Zhang +1

Deployed reinforcement learning agents often face safety requirements that are specified only after training, such as new hazard maps, revised risk thresholds, or behavioral alignm…

cs.LG2026

Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach

Utsav Singh, Souradip Chakraborty, Wesley A. Suttle +6

Hierarchical reinforcement learning (HRL) enables agents to solve complex, long-horizon tasks by decomposing them into manageable sub-tasks. However, HRL methods face two fundament…

cs.LG2026

Improved Sample Complexity For Diffusion Model Training Without Empirical Risk Minimizer Access

Mudit Gaur, Prashant Trivedi, Sasidhar Kunapuli +2

Diffusion models have demonstrated state-of-the-art performance across vision, language, and scientific domains. Despite their empirical success, prior theoretical analyses of the…

cs.LG2026

Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training

Anas Barakat, Souradip Chakraborty, Khushbu Pahwa +1

Pass@k is a widely used performance metric for verifiable large language model tasks, including mathematical reasoning, code generation, and short-answer reasoning. It defines succ…