quantitative finance

Is Deep Hedging Reinforcement Learning?

arXiv:2607.13353

summary

The paper argues that the deep hedging framework, which trains neural network policies via Monte‑Carlo policy‑gradient methods to minimize risk measures, should be classified as reinforcement learning despite lacking intermediate rewards or a value function.

Abstract

The deep hedging framework of Buehler et al. (2019) trains a neural network policy, via Monte Carlo simulation of price paths and stochastic gradient descent, to minimize a risk measure applied to the terminal hedging error. In a recent stream of papers, my coauthors and I have described this technique as reinforcement learning (RL). Several peers have, on occasion, expressed the view that deep hedging does not constitute genuine RL, on two grounds, among others: first, that because feedback is generated only at the terminal date, with no intermediate reward signal, the method cannot constitute genuine RL; and second, that the absence of a value function, a Bellman equation, temporal-difference (TD) learning, and an explicit exploration mechanism disqualifies the method from the RL category altogether, so that it should instead be labeled a neural-network method for stochastic optimal control. The present note argues instead that both objections rest on an unduly narrow, TD-centric reading of what constitutes RL, and that once RL is understood, as it is in the standard references of the field, to include Monte Carlo policy-gradient methods and direct (actor-only) policy search as first-class members, the deep hedging algorithm of Buehler et al. (2019) falls squarely within the RL umbrella.

Topics & keywords

#deep hedging#reinforcement learning#policy gradient#stochastic optimal control#risk measures#neural networksdeep hedgingpolicy gradientMonte Carlo simulationrisk measureactor-only policy searchBellman equation
Is Deep Hedging Reinforcement Learning? · wovepaper