Transfer-Entropy-Regularized Markov Decision Processes
arXiv:1708.09096
Abstract
We consider the framework of transfer-entropy-regularized Markov Decision Process (TERMDP) in which the weighted sum of the classical state-dependent cost and the transfer entropy from the state random process to the control random process is minimized. Although TERMDPs are generally formulated as nonconvex optimization problems, we derive an analytical necessary optimality condition expressed as a finite set of nonlinear equations, based on which an iterative forward-backward computational procedure similar to the Arimoto-Blahut algorithm is proposed. It is shown that every limit point of the sequence generated by the proposed algorithm is a stationary point of the TERMDP. Applications of TERMDPs are discussed in the context of networked control systems theory and non-equilibrium thermodynamics. The proposed algorithm is applied to an information-constrained maze navigation problem, whereby we study how the price of information qualitatively alters the optimal decision polices.
References in corpus (6)
- Second-law-like inequalities with information and their interpretations
- Semidefinite Programming Approach to Gaussian Sequential Rate-Distortion Trade-offs
- A Characterization of the Minimal Average Data Rate that Guarantees a Given Closed-Loop Performance Level
- Rate of Prefix-free Codes in LQG Control Systems
- Information Nonanticipative Rate Distortion Function and Its Applications
- An Upper Bound to Zero-Delay Rate Distortion via Kalman Filtering for Vector Gaussian Sources