activity
20172022
most citedExploration-Exploitation in Constrained MDPs

30 citations · 86 across the 16 of their papers we have counts for

collaborators

24 papers

cs.LG20221 cited

Improved Adaptive Algorithm for Scalable Active Learning with Weak Labeler

Yifang Chen, Karthik Sankararaman, Alessandro Lazaric +6

Active learning with strong and weak labelers considers a practical setting where we have access to both costly but accurate strong labelers and inaccurate but cheap predictions pr…

cs.LG20221 cited

Scalable Representation Learning in Linear Contextual Bandits with Constant Regret Guarantees

Andrea Tirinzoni, Matteo Papini, Ahmed Touati +2

We study the problem of representation learning in stochastic contextual linear bandits. While the primary concern in this domain is usually to find realizable representations (i.e…

cs.LG2022

Reaching Goals is Hard: Settling the Sample Complexity of the Stochastic Shortest Path

Liyu Chen, Andrea Tirinzoni, Matteo Pirotta +1

We study the sample complexity of learning an -optimal policy in the Stochastic Shortest Path (SSP) problem. We first derive sample complexity bounds when the learner has access…

cs.LG20212 cited

Reinforcement Learning in Linear MDPs: Constant Regret and Representation Selection

Matteo Papini, Andrea Tirinzoni, Aldo Pacchiano +3

We study the role of the representation of state-action value functions in regret minimization in finite-horizon Markov Decision Processes (MDPs) with linear structure. We first de…

cs.LG20212 cited

A Fully Problem-Dependent Regret Lower Bound for Finite-Horizon MDPs

Andrea Tirinzoni, Matteo Pirotta, Alessandro Lazaric

We derive a novel asymptotic problem-dependent lower-bound for regret minimization in finite-horizon tabular Markov Decision Processes (MDPs). While, similar to prior work (e.g., f…

cs.LG20218 cited

Leveraging Good Representations in Linear Contextual Bandits

Matteo Papini, Andrea Tirinzoni, Marcello Restelli +2

The linear contextual bandit literature is mostly focused on the design of efficient learning algorithms for a given representation. However, a contextual bandit problem may admit…