11 papers · 1 filter
Provably Efficient Off-Policy Adversarial Imitation Learning with Convergence Guarantees
Yilei Chen, Vittorio Giammarino, James Queeney +1
Adversarial Imitation Learning (AIL) faces challenges with sample inefficiency because of its reliance on sufficient on-policy data to evaluate the performance of the current polic…
Multiple-policy Evaluation via Density Estimation
Yilei Chen, Aldo Pacchiano, Ioannis Ch. Paschalidis
We study the multiple-policy evaluation problem where we are given a set of policies and the goal is to evaluate their performance (expected total reward over a fixed horizon)…
Geometric Re-Analysis of Classical MDP Solving Algorithms
Arsenii Mustafin, Aleksei Pakharev, Alex Olshevsky +1
We build on a recently introduced geometric interpretation of Markov Decision Processes (MDPs) to analyze classical MDP-solving algorithms: Value Iteration (VI) and Policy Iteratio…
MDP Geometry, Normalization and Reward Balancing Solvers
Arsenii Mustafin, Aleksei Pakharev, Alex Olshevsky +1
We present a new geometric interpretation of Markov Decision Processes (MDPs) with a natural normalization procedure that allows us to adjust the value function at each state witho…
Analysis of Value Iteration Through Absolute Probability Sequences
Arsenii Mustafin, Sebastien Colla, Alex Olshevsky +1
Value Iteration is a widely used algorithm for solving Markov Decision Processes (MDPs). While previous studies have extensively analyzed its convergence properties, they primarily…
Generalized Policy Improvement Algorithms with Theoretically Supported Sample Reuse
James Queeney, Ioannis Ch. Paschalidis, Christos G. Cassandras
We develop a new class of model-free deep reinforcement learning algorithms for data-driven, learning-based control. Our Generalized Policy Improvement algorithms combine the polic…