activity
20102015
most citedApprenticeship Learning using Inverse Reinforcement Learning and Gradient Methods

157 citations · 377 across the 6 of their papers we have counts for

collaborators

6 papers

cs.LG2015108 cited

Cascading Bandits: Learning to Rank in the Cascade Model

Branislav Kveton, Csaba Szepesvari, Zheng Wen +1

A search engine usually outputs a list of web pages. The user examines this list, from the first web page to the last, and chooses the first attractive page. This model of user…

cs.AI201413 cited

On Minimax Optimal Offline Policy Evaluation

Lihong Li, Remi Munos, Csaba Szepesvari

This paper studies the off-policy evaluation problem, where one aims to estimate the value of a target policy based on a sample of observations collected by another policy. We firs…

cs.LG2012157 cited

Apprenticeship Learning using Inverse Reinforcement Learning and Gradient Methods

Gergely Neu, Csaba Szepesvari

In this paper we propose a novel gradient algorithm to learn a policy from an expert's observed behavior assuming that the expert behaves optimally with respect to some unknown rew…

cs.AI20123 cited

Speeding Up Planning in Markov Decision Processes via Automatically Constructed Abstractions

Alejandro Isaza, Csaba Szepesvari, Vadim Bulitko +1

In this paper, we consider planning in stochastic shortest path (SSP) problems, a subclass of Markov Decision Problems (MDP). We focus on medium-size problems whose state space can…

cs.LG20125 cited

PAC-Bayesian Policy Evaluation for Reinforcement Learning

Mahdi MIlani Fard, Joelle Pineau, Csaba Szepesvari

Bayesian priors offer a compact yet general means of incorporating domain knowledge into many learning tasks. The correctness of the Bayesian analysis and inference, however, large…

stat.ML201091 cited

Estimation of Rényi Entropy and Mutual Information Based on Generalized Nearest-Neighbor Graphs

Dávid Pál, Barnabás Póczos, Csaba Szepesvári

We present simple and computationally efficient nonparametric estimators of Rényi entropy and mutual information based on an i.i.d. sample drawn from an unknown, absolutely continu…