activity
20242026
collaborators

5 papers

cs.LG2026

Replicable Bandits with UCB based Exploration

Rohan Deb, Udaya Ghai, Karan Singh +1

We study replicable algorithms for stochastic multi-armed bandits (MAB) and linear bandits with UCB (Upper Confidence Bound) based exploration. A bandit algorithm is -replicabl…

cs.LG2026

Inference Time Policy Optimization for Offline RL with Differentiable World Models

Rohan Deb, Stephen J. Wright, Arindam Banerjee

Offline Reinforcement Learning (RL) learns optimal policies from fixed datasets, training a policy once and deploying it at inference time without further refinement. Inspired by m…

cs.LG2026

Plan Before You Trade: Inference-Time Optimization for RL Trading Agents

Eun Go, Rohan Deb, Arindam Banerjee

Reinforcement learning agents for portfolio management are typically trained and deployed as static policies, with no mechanism for using price forecasts at inference time. We prop…

cs.LG2025

Beyond Johnson-Lindenstrauss: Uniform Bounds for Sketched Bilinear Forms

Rohan Deb, Qiaobo Li, Mayank Shrivastava +1

Uniform bounds on sketched inner products of vectors or matrices underpin several important computational and statistical results in machine learning and randomized algorithms, inc…

cs.LG2024

Conservative Contextual Bandits: Beyond Linear Representations

Rohan Deb, Mohammad Ghavamzadeh, Arindam Banerjee

Conservative Contextual Bandits (CCBs) address safety in sequential decision making by requiring that an agent's policy, along with minimizing regret, also satisfies a safety const…