papers

Publications (13)

cs.LG2025

Beyond Johnson-Lindenstrauss: Uniform Bounds for Sketched Bilinear Forms

Rohan Deb, Qiaobo Li, Mayank Shrivastava +1

Uniform bounds on sketched inner products of vectors or matrices underpin several important computational and statistical results in machine learning and randomized algorithms, inc…

cs.LG2026

KMM-CP: Practical Conformal Prediction under Covariate Shift via Selective Kernel Mean Matching

Siddhartha Laghuvarapu, Rohan Deb, Jimeng Sun

Uncertainty quantification is essential for deploying machine learning models in high-stakes domains such as scientific discovery and healthcare. Conformal Prediction (CP) provides…

cs.LG2025

FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain

Rohan Deb, Kiran Thekumparampil, Kousha Kalantari +3

Supervised fine-tuning (SFT) is a standard approach to adapting large language models (LLMs) to new domains. In this work, we improve the statistical efficiency of SFT by selecting…

cs.LG2021

Gradient Temporal Difference with Momentum: Stability and Convergence

Rohan Deb, Shalabh Bhatnagar

Gradient temporal difference (Gradient TD) algorithms are a popular class of stochastic approximation (SA) algorithms used for policy evaluation in reinforcement learning. Here, we…

cs.LG2021

Schedule Based Temporal Difference Algorithms

Rohan Deb, Meet Gandhi, Shalabh Bhatnagar

Learning the value function of a given policy from data samples is an important problem in Reinforcement Learning. TD() is a popular class of algorithms to solve this problem.…

cs.LG2022

Does Momentum Help? A Sample Complexity Analysis

Swetha Ganesh, Rohan Deb, Gugan Thoppe +1

Stochastic Heavy Ball (SHB) and Nesterov's Accelerated Stochastic Gradient (ASG) are popular momentum methods in stochastic optimization. While benefits of such acceleration ideas…

cs.LG2026

Replicable Bandits with UCB based Exploration

Rohan Deb, Udaya Ghai, Karan Singh +1

We study replicable algorithms for stochastic multi-armed bandits (MAB) and linear bandits with UCB (Upper Confidence Bound) based exploration. A bandit algorithm is -replicabl…

cs.LG2026

Inference Time Policy Optimization for Offline RL with Differentiable World Models

Rohan Deb, Stephen J. Wright, Arindam Banerjee

Offline Reinforcement Learning (RL) learns optimal policies from fixed datasets, training a policy once and deploying it at inference time without further refinement. Inspired by m…

cs.LG2026

Plan Before You Trade: Inference-Time Optimization for RL Trading Agents

Eun Go, Rohan Deb, Arindam Banerjee

Reinforcement learning agents for portfolio management are typically trained and deployed as static policies, with no mechanism for using price forecasts at inference time. We prop…

cs.LG2024

Conservative Contextual Bandits: Beyond Linear Representations

Rohan Deb, Mohammad Ghavamzadeh, Arindam Banerjee

Conservative Contextual Bandits (CCBs) address safety in sequential decision making by requiring that an agent's policy, along with minimizing regret, also satisfies a safety const…

cs.LG2023

Think Before You Duel: Understanding Complexities of Preference Learning under Constrained Resources

Rohan Deb, Aadirupa Saha

We consider the problem of reward maximization in the dueling bandit setup along with constraints on resource consumption. As in the classic dueling bandits, at each round the lear…

cs.LG2023

Contextual Bandits with Online Neural Regression

Rohan Deb, Yikun Ban, Shiliang Zuo +2

Recent works have shown a reduction from contextual bandits to online regression under a realizability assumption [Foster and Rakhlin, 2020, Foster and Krishnamurthy, 2021]. In thi…

eess.SY2025

Multi Timescale Stochastic Approximation: Stability and Convergence

Rohan Deb, Swetha Ganesh, Shalabh Bhatnagar

This paper presents the first sufficient conditions that guarantee the stability and almost sure convergence of multi-timescale stochastic approximation (SA) iterates. It extends t…