Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Performative Policy Gradient: Optimality in Performative Reinforcement Learning
Debabrota Basu, Udvas Das, Brahim Driss +1
Post-deployment machine learning algorithms often influence the environments they act in, and thus shift the underlying dynamics that the standard reinforcement learning (RL) metho…
cs.LG2025
StaQ: a Finite Memory Approach to Discrete Action Policy Mirror Descent
Alena Shilova, Alex Davey, Brahim Driss +1
In Reinforcement Learning (RL), regularization with a Kullback-Leibler divergence that penalizes large deviations between successive policies has emerged as a popular tool both in…