3 papers
cs.LG2026
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information
Priyank Agrawal, Ankur Samanta, Shervin Ghasemlou +5
Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot gene…
cs.LG2025
Q-learning with Posterior Sampling
Priyank Agrawal, Shipra Agrawal, Azmat Azati
Bayesian posterior sampling techniques have demonstrated superior empirical performance in many exploration-exploitation settings. However, their theoretical analysis remains a cha…
cs.LG2024
On the Convergence of Single-Timescale Actor-Critic
Navdeep Kumar, Priyank Agrawal, Giorgia Ramponi +2
We analyze the global convergence of the single-timescale actor-critic (AC) algorithm for the infinite-horizon discounted Markov Decision Processes (MDPs) with finite state spaces.…