4 papers
Direct Advantage Estimation for Scalable and Sample-efficient Deep Reinforcement Learning
Hsiao-Ru Pan, Bernhard Schölkopf
Direct Advantage Estimation (DAE) has been shown to improve the sample efficiency of deep reinforcement learning algorithms. However, its reliance on full environment observability…
On the Variance of Temporal Difference Learning and its Reduction Using Control Variates
Hsiao-Ru Pan, Bernhard Schölkopf
We analyze the variance of temporal difference (TD) learning using the phased setting with tabular representation, and show that one of the mechanisms behind its ability to reduce…
Intrinsically Interpretable Attention via Sparse Post-Training
Florent Draye, Anson Lei, Hsiao-Ru Pan +2
We introduce a simple post-training method that makes transformer attention sparse without sacrificing performance. Applying a flexible sparsity regularisation under a constrained-…
Homomorphism Autoencoder -- Learning Group Structured Representations from Observed Transitions
Hamza Keurti, Hsiao-Ru Pan, Michel Besserve +2
How can agents learn internal models that veridically represent interactions with the real world is a largely open question. As machine learning is moving towards representations c…