29 papers
Beyond Marginal Validity: Finite-Sample Guarantees for Localized Conformal Prediction
Anton Conrad, Rustam Isaev, Denis Belomestny +2
Conformal prediction endows arbitrary black-box predictors with finite-sample, distribution-free marginal coverage, yet marginal validity can hide severe covariate-specific miscali…
Adaptive Cumulative Mass Calibration with Conformal Prediction
Daniil Kazantsev, Eric Moulines, Maxim Panov +2
Reliable probability estimates by classifiers are essential in high-risk applications. In practice, however, predicted probabilities are often miscalibrated, and many existing post…
Gaussian Approximation and Multiplier Bootstrap for Stochastic Gradient Descent
Marina Sheshukova, Sergey Samsonov, Denis Belomestny +4
In this paper, we establish the non-asymptotic validity of the multiplier bootstrap procedure for constructing the confidence sets using the Stochastic Gradient Descent (SGD) algor…
Beyond Softmax and Entropy: Convergence Rates of Policy Gradients with f-SoftArgmax Parameterization & Coupled Regularization
Safwan Labbi, Daniil Tiapkin, Paul Mangold +1
Policy gradient methods are known to be highly sensitive to the choice of policy parameterization. In particular, the widely used softmax parameterization can induce ill-conditione…
On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments
Safwan Labbi, Paul Mangold, Daniil Tiapkin +1
We provide global convergence rates for vanilla and entropy-regularized federated softmax stochastic policy gradient (FedPG) with local training. We show that FedPG converges to a…
Proximal Point Nash Learning from Human Feedback
Daniil Tiapkin, Daniele Calandriello, Denis Belomestny +5
Traditional Reinforcement Learning from Human Feedback (RLHF) often relies on reward models, frequently assuming preference structures like the Bradley--Terry model, which may not…