3 papers
cs.LG2024
Data-Driven Upper Confidence Bounds with Near-Optimal Regret for Heavy-Tailed Bandits
Ambrus Tamás, Szabolcs Szentpéteri, Balázs Csanád Csáji
Stochastic multi-armed bandits (MABs) provide a fundamental reinforcement learning model to study sequential decision making in uncertain environments. The upper confidence bounds…
eess.SY2024
Finite-Sample Identification of Linear Regression Models with Residual-Permuted Sums
Szabolcs Szentpéteri, Balázs Csanád Csáji
This letter studies a distribution-free, finite-sample data perturbation (DP) method, the Residual-Permuted Sums (RPS), which is an alternative of the Sign-Perturbed Sums (SPS) alg…
eess.SY2024
Non-Asymptotic State-Space Identification of Closed-Loop Stochastic Linear Systems using Instrumental Variables
Szabolcs Szentpéteri, Balázs Csanád Csáji
The paper suggests a generalization of the Sign-Perturbed Sums (SPS) finite sample system identification method for the identification of closed-loop observable stochastic linear s…