Empirical Measure Large Deviations for Reinforced Chains on Finite Spaces
arXiv:2205.09291
Abstract
Let be a transition probability kernel on a finite state space such that for all . Consider a reinforced chain given as a sequence of -valued random variables, defined recursively according to, We establish a large deviation principle for . The rate function takes a strikingly different form than the Donsker-Varadhan rate function associated with the empirical measure of the Markov chain with transition kernel and is described in terms of a novel deterministic infinite horizon discounted cost control problem with an associated linear controlled dynamics and a nonlinear running cost involving the relative entropy function. Proofs are based on an analysis of time-reversal of controlled dynamics in representations for log-transforms of exponential moments, and on weak convergence methods.