activity
20152026
most citedGeneralized Proximal Policy Optimization with Sample Reuse

21 citations · 59 across the 19 of their papers we have counts for

collaborators
Showing 2020Show all

8 papers · 1 filter

cs.LG2020

Uncertainty-Aware Policy Optimization: A Robust, Adaptive Trust Region Approach

James Queeney, Ioannis Ch. Paschalidis, Christos G. Cassandras

In order for reinforcement learning techniques to be useful in real-world decision making processes, they must be able to produce robust performance from limited data. Deep policy…

cs.LG2020

Provable Hierarchical Imitation Learning via EM

Zhiyu Zhang, Ioannis Paschalidis

Due to recent empirical successes, the options framework for hierarchical reinforcement learning is gaining increasing popularity. Rather than learning from rewards which suffers f…

cs.AI2020

Explainability of Intelligent Transportation Systems using Knowledge Compilation: a Traffic Light Controller Case

Salomón Wollenstein-Betech, Christian Muise, Christos G. Cassandras +2

Usage of automated controllers which make decisions on an environment are widespread and are often based on black-box models. We use Knowledge Compilation theory to bring explainab…

stat.ML2020

Robust Grouped Variable Selection Using Distributionally Robust Optimization

Ruidi Chen, Ioannis Ch. Paschalidis

We propose a Distributionally Robust Optimization (DRO) formulation with a Wasserstein-based uncertainty set for selecting grouped variables under perturbations on the data for bot…

stat.ML20204 cited

Robustified Multivariate Regression and Classification Using Distributionally Robust Optimization under the Wasserstein Metric

Ruidi Chen, Ioannis Ch. Paschalidis

We develop Distributionally Robust Optimization (DRO) formulations for Multivariate Linear Regression (MLR) and Multiclass Logistic Regression (MLG) when both the covariates and re…

math.OC202012 cited

Local SGD With a Communication Overhead Depending Only on the Number of Workers

Artin Spiridonoff, Alex Olshevsky, Ioannis Ch. Paschalidis

We consider speeding up stochastic gradient descent (SGD) by parallelizing it across multiple workers. We assume the same data set is shared among workers, who can take SGD ste…