2 papers
cs.LG2025
Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function Approximation
Fengdi Che, Chenjun Xiao, Jincheng Mei +6
We prove that the combination of a target network and over-parameterized linear function approximation establishes a weaker convergence condition for bootstrapped value estimation…
cs.LG2025
Average-DICE: Stationary Distribution Correction by Regression
Fengdi Che, Bryan Chan, Chen Ma +1
Off-policy policy evaluation (OPE), an essential component of reinforcement learning, has long suffered from stationary state distribution mismatch, undermining both stability and…