most citedStatistical Inference for Policy Evaluation with Temporal Difference Learning

1 citations · 1 across the 4 of their papers we have counts for

collaborators

5 papers

math.ST2026

Prediction-Only Distillation in Linear and Logistic Regression

Hien Dang, Pratik Patil, Alessandro Rinaldo

Self-distillation (SD) is typically studied when the student is retrained on the teacher's original training inputs. In many practical deployments, however, the labeled training da…

stat.ML20261 cited

Statistical Inference for Policy Evaluation with Temporal Difference Learning

Weichen Wu, Gen Li, Yuting Wei +1

We investigate the statistical properties of Temporal Difference (TD) learning with Polyak-Ruppert averaging, arguably one of the most widely used algorithms in reinforcement learn…

stat.ML2026

Uncertainty quantification for Markov chain induced martingales with application to temporal difference learning

Weichen Wu, Yuting Wei, Alessandro Rinaldo

We establish novel and general high-dimensional concentration inequalities and Berry-Esseen bounds for vector-valued martingales induced by Markov chains. We apply these results to…

math.PR2026

Berry-Esseen bounds for multivariate martingale difference sequences in the Kolmogorov distance

Weichen Wu, Dung Le, Arun Kumar Kuchibhotla +1

We derive new Gaussian approximations for finite martingale difference sequences in with respect to the Kolmogorov distance. Under appropriate conditions, our bounds…

math.ST2026

Optimal Unconstrained Self-Distillation in Ridge Regression: Strict Improvements, Precise Asymptotics, and One-Shot Tuning

Hien Dang, Pratik Patil, Alessandro Rinaldo

Self-distillation (SD) is the process of retraining a student on a mixture of ground-truth labels and the teacher's own predictions using the same architecture and training data. A…