paper

Self-Consistent Adjoint Policy Iteration for Constrained Dynamic Portfolio Choice

arXiv:2608.17808

Abstract

We develop simulation-based policy iteration for continuous-time portfolio choice with predictable returns and convex constraints. Each outer step re-evaluates a fixed-latent open-loop backpropagation-through-time (OL-BPTT) adjoint after deployment and solves the constrained update. Shifted-adjoint cancellation controls the adjoint--HJB Hamiltonian-gradient discrepancy by the policy-improvement residual. For CRRA portfolios, exact HJB policy iteration identifies the optimal reduced value factor, while population OL-BPTT iteration converges globally when the adjoint update is directionally improving and approximate stationarity is asymptotically HJB-compatible. A theorem-matched occupation audit yields maximal upper endpoints of for the primitive directional ratio and for a stronger norm-relative ratio, both against the half-step threshold . In the high-precision -- occupancy/broad-anchor design of a three-factor, fifty-asset benchmark, current-policy re-evaluation outperforms matched pooled refinement under the on-policy and broad evaluation laws.

Self-Consistent Adjoint Policy Iteration for Constrained Dynamic Portfolio Choice · wovepaper