#policy gradient

4 results
cs.LG2026

Policy Gradient Steering: Interventions from Behavioral Objectives

Yoann Poupart, Aurélie Beynier, Nicolas Maudet

The paper introduces Policy Gradient Steering (PGS), a reinforcement‑learning based method that uses temporary behavioral objectives to compute removable task vectors for steering…

#reinforcement learning#policy gradient#behavioral steering#composable interventions
q-fin.CP2026

Is Deep Hedging Reinforcement Learning?

Frédéric Godin

The paper argues that the deep hedging framework, which trains neural network policies via Monte‑Carlo policy‑gradient methods to minimize risk measures, should be classified as re…

#deep hedging#reinforcement learning#policy gradient#stochastic optimal control
cs.RO2026

TADPO: Reinforcement Learning Goes Off-road

Zhouchonghao Wu, Raymond Song, Vedant Mundheda +3

The paper introduces TADPO, a policy‑gradient method that extends PPO with teacher‑student guidance, and uses it in a vision‑based end‑to‑end reinforcement learning system for high…

#off-road navigation#reinforcement learning#policy gradient#sim-to-real transfer
cs.LG2026

Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents

Tianyu Ding, Jianhong Xin, Juan Pablo De la Cruz Weinstein

The paper introduces Sibling-Guided Credit Distillation (SGCD), a technique that uses self‑distillation to assign finer‑grained credit to actions in long‑horizon tool‑use reinforce…

#long-horizon reinforcement learning#tool-use agents#credit distillation#policy gradient