activity
20152026
most citedGeneralized Proximal Policy Optimization with Sample Reuse

21 citations · 59 across the 20 of their papers we have counts for

collaborators
Showing 2024Show all

10 papers · 1 filter

eess.SY2024

Network-Based Epidemic Control Through Optimal Travel and Quarantine Management

Mahtab Talaei, Apostolos I. Rikos, Alex Olshevsky +2

Motivated by the swift global transmission of infectious diseases, we present a comprehensive framework for network-based epidemic control. Our aim is to curb epidemics using two d…

cs.LG2024

Visually Robust Adversarial Imitation Learning from Videos with Contrastive Learning

Vittorio Giammarino, James Queeney, Ioannis Ch. Paschalidis

We propose C-LAIfO, a computationally efficient algorithm designed for imitation learning from videos in the presence of visual mismatch between agent and expert domains. We analyz…

cs.LG2024

MDP Geometry, Normalization and Reward Balancing Solvers

Arsenii Mustafin, Aleksei Pakharev, Alex Olshevsky +1

We present a new geometric interpretation of Markov Decision Processes (MDPs) with a natural normalization procedure that allows us to adjust the value function at each state witho…

cs.LG2024

On Value Iteration Convergence in Connected MDPs

Arsenii Mustafin, Alex Olshevsky, Ioannis Ch. Paschalidis

This paper establishes that an MDP with a unique optimal policy and ergodic associated transition matrix ensures the convergence of various versions of the Value Iteration algorith…

cs.LG2024

Provably Efficient Off-Policy Adversarial Imitation Learning with Convergence Guarantees

Yilei Chen, Vittorio Giammarino, James Queeney +1

Adversarial Imitation Learning (AIL) faces challenges with sample inefficiency because of its reliance on sufficient on-policy data to evaluate the performance of the current polic…

cs.LG2024

Multiple-policy Evaluation via Density Estimation

Yilei Chen, Aldo Pacchiano, Ioannis Ch. Paschalidis

We study the multiple-policy evaluation problem where we are given a set of policies and the goal is to evaluate their performance (expected total reward over a fixed horizon)…