2 papers
cs.LG2026
PROMA: Projected Microbatch Accumulation for Reference-Free Proximal Policy Updates
Nilin Abrahamsen
This note introduces Projected Microbatch Accumulation (PROMA), a reference-free proximal policy method that controls KL divergence by projecting away high-variance components of t…
cs.LG2025
ISOPO: Proximal policy gradients without pi-old
Nilin Abrahamsen
This note introduces Isometric Policy Optimization (ISOPO), an efficient method to approximate the natural policy gradient in a single gradient step. In comparison, existing proxim…