7 citations · 12 across the 3 of their papers we have counts for
4 papers
Tsetlin Machine for Solving Contextual Bandit Problems
Raihan Seraj, Jivitesh Sharma, Ole-Christoffer Granmo
This paper introduces an interpretable contextual bandit algorithm using Tsetlin Machines, which solves complex pattern recognition tasks using propositional logic. The proposed ba…
Approximate information state for approximate planning and reinforcement learning in partially observed systems
Jayakumar Subramanian, Amit Sinha, Raihan Seraj +1
We propose a theoretical framework for approximate planning and learning in partially observed systems. Our framework is based on the fundamental notion of information state. We pr…
Doubly Robust Off-Policy Actor-Critic Algorithms for Reinforcement Learning
Riashat Islam, Raihan Seraj, Samin Yeasar Arnob +1
We study the problem of off-policy critic evaluation in several variants of value-based off-policy actor-critic algorithms. Off-policy actor-critic algorithms require an off-policy…
Entropy Regularization with Discounted Future State Distribution in Policy Gradient Methods
Riashat Islam, Raihan Seraj, Pierre-Luc Bacon +1
The policy gradient theorem is defined based on an objective with respect to the initial distribution over states. In the discounted case, this results in policies that are optimal…