Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
KL-Regularized Reinforcement Learning is Designed to Mode Collapse
Anthony GX-Chen, Jatin Prakash, Jeff Guo +2
It is commonly believed that optimizing the reverse KL divergence results in "mode seeking", while optimizing forward KL results in "mass covering", with the latter being preferred…
cs.LG2022
A Generalized Bootstrap Target for Value-Learning, Efficiently Combining Value and Feature Predictions
Anthony GX-Chen, Veronica Chelu, Blake A. Richards +1
Estimating value functions is a core component of reinforcement learning algorithms. Temporal difference (TD) learning algorithms use bootstrapping, i.e. they update the value func…