#policy optimization
21 papers · 1 filter
-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
Jiawei Xu, Minghui Liu, Juzheng Zhang +2
The paper proposes β‑OPSD, a generalized on‑policy self‑distillation method that treats the KL regularization weight as a tunable parameter, enabling a controlled interpolation bet…
APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems
Shentong Mo, Yatao Bian
The paper introduces Atomic Policy Optimization (APO), an unsupervised method that learns to predict 3D structures of atomic systems by optimizing a policy with dual rewards for st…
TAPO: Transition-Aware Policy Optimization for LLM Agents
Cong Li, Peixi Peng, Yisen Zhao +4
The paper introduces TAPO, a training framework that augments reinforcement learning for large language model agents with action‑conditioned next‑observation prediction, improving…
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts
Ken Ding
The paper introduces LoRA Scaffolded Policy Optimization (LSPO), a sampling-time low-rank adapter method that recovers gradient information for reinforcement learning on difficult…
Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold
Songshuo Lu, Zhi Chen, Yaohua Tang
The paper proposes an expand‑then‑compress framework that builds a diverse set of RL‑trained teacher models and then distills them into a single student model, improving reasoning,…
ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning
Zhenrong Zhang, Fei Wu, Jun Du +2
The paper presents ReDiPPO, a PPO-based reinforcement learning framework that leverages reference answers to guide value estimation and reweights token-level advantages based on di…