Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
PROMA: Projected Microbatch Accumulation for Reference-Free Proximal Policy Updates
Nilin Abrahamsen
This note introduces Projected Microbatch Accumulation (PROMA), a reference-free proximal policy method that controls KL divergence by projecting away high-variance components of t…
cs.LG2025
ISOPO: Proximal policy gradients without pi-old
Nilin Abrahamsen
This note introduces Isometric Policy Optimization (ISOPO), an efficient method to approximate the natural policy gradient in a single gradient step. In comparison, existing proxim…
cs.LG2024
Convergence of variational Monte Carlo simulation and scale-invariant pre-training
Nilin Abrahamsen, Zhiyan Ding, Gil Goldshlager +1
We provide theoretical convergence bounds for the variational Monte Carlo (VMC) method as applied to optimize neural network wave functions for the electronic structure problem. We…