paper

Tight Bounds for Linear and Non-Linear Contraction of Divergences via Duality

arXiv:2402.11200

Abstract

We develop a novel framework for bounding the contraction of information divergences, using duality and associated norms in Orlicz spaces. By working in the dual space, we obtain a principled approach to bounding both distribution-dependent strong data-processing inequality (SDPI) constants and \(F_φ\)-curves of divergences. Our bounds are either available in closed form or reducible to one-dimensional convex optimisation problems, in contrast to the infinite-dimensional optimisation problems that characterise SDPIs. These bounds depend on the densities of the reverse kernels with respect to a reference measure. To the best of our knowledge, they are the first universal closed-form bounds on distribution-dependent SDPI constants. We establish tightness for the \(χ^2\)-divergence on several important channel classes, including full-rank binary kernels. We apply our results to several settings. In particular, we derive bounds on the mixing times of Markov chains, including chains with heavy-tailed stationary distributions; obtain improved bounds on burn-in periods for Markov chain Monte Carlo; and strengthen concentration-of-measure bounds for dependent random variables.

An old version of the work was accepted for presentation at the Conference on Learning Theory (COLT) 2024 The new version refocuses the work on linear and non-linear contraction of divergences