#policy optimization

topicpolicy optimization

21 papers · 1 filter

cs.LG2026

-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

Jiawei Xu, Minghui Liu, Juzheng Zhang +2

The paper proposes β‑OPSD, a generalized on‑policy self‑distillation method that treats the KL regularization weight as a tunable parameter, enabling a controlled interpolation bet…

cs.LG2026

APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems

Shentong Mo, Yatao Bian

The paper introduces Atomic Policy Optimization (APO), an unsupervised method that learns to predict 3D structures of atomic systems by optimizing a policy with dual rewards for st…

cs.LG2026

TAPO: Transition-Aware Policy Optimization for LLM Agents

Cong Li, Peixi Peng, Yisen Zhao +4

The paper introduces TAPO, a training framework that augments reinforcement learning for large language model agents with action‑conditioned next‑observation prediction, improving…

cs.LG2026

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts

Ken Ding

The paper introduces LoRA Scaffolded Policy Optimization (LSPO), a sampling-time low-rank adapter method that recovers gradient information for reinforcement learning on difficult…

cs.LG2026

Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold

Songshuo Lu, Zhi Chen, Yaohua Tang

The paper proposes an expand‑then‑compress framework that builds a diverse set of RL‑trained teacher models and then distills them into a single student model, improving reasoning,…

cs.AI2026

ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning

Zhenrong Zhang, Fei Wu, Jun Du +2

The paper presents ReDiPPO, a PPO-based reinforcement learning framework that leverages reference answers to guide value estimation and reweights token-level advantages based on di…