Monotone Neural Policy Iteration for High-Dimensional First-Order Hamilton--Jacobi--Bellman Equations
arXiv:2605.07116
Abstract
We analyze a neural semi-discrete method for high-dimensional first-order Hamilton-Jacobi-Bellman (HJB) equations with known or learned dynamics. Centered differences and an artificial viscosity define a monotone operator evaluated through shifted network queries; policy iteration solves the resulting Bellman equation without a tensor grid. At fixed , the sharp componentwise condition turns every frozen-policy operator into a nearest-neighbor Markov-chain generator with a policy-independent total jump rate. Uniformization gives whole-space well-posedness for measurable feedbacks, an explicit Poisson-tail bound on the numerical domain of dependence, and boundary-free localization. The representation also yields a posteriori policy-evaluation bounds that account for residual and learned-model errors. A greedy-gap analysis controls inexact policy iteration at fixed ; a separate consistency estimate connects the semi-discrete equation to the continuous HJB equation. Experiments reproduce the extremal tail, show rates consistent with and nearly -independent exact-policy-iteration decay, and assess empirical estimator effectivity. A nonsmooth example shows that the continuous residual can miss a non-viscosity solution, whereas the shifted residual detects the defect. Further tests provide a structured interval-verified certificate calibration, an early-budget benefit of policy freezing for bang-bang control, and learned-dynamics diagnostics. A structured nonlinear problem with active compact-control constraints is tested against a manufactured semi-discrete reference through .