4 papers
Regret and Sample Complexity of Online Q-Learning via Concentration of Stochastic Approximation with Time-Inhomogeneous Markov Chains
Rahul Singh, Siddharth Chandak, Eric Moulines +2
We present the first regret bound for classical online Q-learning in infinite-horizon discounted Markov decision processes (MDPs), without relying on optimism or bonus terms. We fi…
Robustness Verification of Binary Neural Networks: An Ising and Quantum-Inspired Framework
Rahul Singh, Seyran Saeedi, Zheng Zhang
Binary neural networks (BNNs) are increasingly deployed in edge computing applications due to their low hardware complexity and high energy efficiency. However, verifying the robus…
Policy Zooming: Adaptive Discretization-based Infinite-Horizon Average-Reward Reinforcement Learning
Avik Kar, Rahul Singh
We study the infinite-horizon average-reward reinforcement learning (RL) for continuous space Lipschitz MDPs in which an agent can play policies from a given set . The proposed…
Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces
Avik Kar, Rahul Singh
We study infinite-horizon average-reward reinforcement learning (RL) for Lipschitz MDPs, a broad class that subsumes several important classes such as linear and RKHS MDPs, functio…