activity
20182022
most citedLeveraging Good Representations in Linear Contextual Bandits

8 citations · 21 across the 9 of their papers we have counts for

collaborators

13 papers

cs.LG20221 cited

Scalable Representation Learning in Linear Contextual Bandits with Constant Regret Guarantees

Andrea Tirinzoni, Matteo Papini, Ahmed Touati +2

We study the problem of representation learning in stochastic contextual linear bandits. While the primary concern in this domain is usually to find realizable representations (i.e…

cs.LG2022

Reaching Goals is Hard: Settling the Sample Complexity of the Stochastic Shortest Path

Liyu Chen, Andrea Tirinzoni, Matteo Pirotta +1

We study the sample complexity of learning an -optimal policy in the Stochastic Shortest Path (SSP) problem. We first derive sample complexity bounds when the learner has access…

cs.AI20211 cited

Dealing With Misspecification In Fixed-Confidence Linear Top-m Identification

Clémence Réda, Andrea Tirinzoni, Rémy Degenne

We study the problem of the identification of m arms with largest means under a fixed error rate (fixed-confidence Top-m identification), for misspecified linear bandit models.…

cs.LG20212 cited

Reinforcement Learning in Linear MDPs: Constant Regret and Representation Selection

Matteo Papini, Andrea Tirinzoni, Aldo Pacchiano +3

We study the role of the representation of state-action value functions in regret minimization in finite-horizon Markov Decision Processes (MDPs) with linear structure. We first de…

cs.LG20212 cited

A Fully Problem-Dependent Regret Lower Bound for Finite-Horizon MDPs

Andrea Tirinzoni, Matteo Pirotta, Alessandro Lazaric

We derive a novel asymptotic problem-dependent lower-bound for regret minimization in finite-horizon tabular Markov Decision Processes (MDPs). While, similar to prior work (e.g., f…

cs.LG2021

Meta-Reinforcement Learning by Tracking Task Non-stationarity

Riccardo Poiani, Andrea Tirinzoni, Marcello Restelli

Many real-world domains are subject to a structured non-stationarity which affects the agent's goals and the environmental dynamics. Meta-reinforcement learning (RL) has been shown…