3 papers
cs.LG2026
Direct Advantage Estimation for Scalable and Sample-efficient Deep Reinforcement Learning
Hsiao-Ru Pan, Bernhard Schölkopf
Direct Advantage Estimation (DAE) has been shown to improve the sample efficiency of deep reinforcement learning algorithms. However, its reliance on full environment observability…
cs.LG2026
On the Variance of Temporal Difference Learning and its Reduction Using Control Variates
Hsiao-Ru Pan, Bernhard Schölkopf
We analyze the variance of temporal difference (TD) learning using the phased setting with tabular representation, and show that one of the mechanisms behind its ability to reduce…
cs.LG2026
Intrinsically Interpretable Attention via Sparse Post-Training
Florent Draye, Anson Lei, Hsiao-Ru Pan +2
We introduce a simple post-training method that makes transformer attention sparse without sacrificing performance. Applying a flexible sparsity regularisation under a constrained-…