#policy gradient
4 resultsPolicy Gradient Steering: Interventions from Behavioral Objectives
Yoann Poupart, Aurélie Beynier, Nicolas Maudet
The paper introduces Policy Gradient Steering (PGS), a reinforcement‑learning based method that uses temporary behavioral objectives to compute removable task vectors for steering…
Is Deep Hedging Reinforcement Learning?
Frédéric Godin
The paper argues that the deep hedging framework, which trains neural network policies via Monte‑Carlo policy‑gradient methods to minimize risk measures, should be classified as re…
TADPO: Reinforcement Learning Goes Off-road
Zhouchonghao Wu, Raymond Song, Vedant Mundheda +3
The paper introduces TADPO, a policy‑gradient method that extends PPO with teacher‑student guidance, and uses it in a vision‑based end‑to‑end reinforcement learning system for high…
Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents
Tianyu Ding, Jianhong Xin, Juan Pablo De la Cruz Weinstein
The paper introduces Sibling-Guided Credit Distillation (SGCD), a technique that uses self‑distillation to assign finer‑grained credit to actions in long‑horizon tool‑use reinforce…