16 citations · 27 across the 2 of their papers we have counts for
6 papers
Normalized Attention Without Probability Cage
Oliver Richter, Roger Wattenhofer
Attention architectures are widely used; they recently gained renewed popularity with Transformers yielding a streak of state of the art results. Yet, the geometrical implications…
On Identifiability in Transformers
Gino Brunner, Yang Liu, Damián Pascual +3
In this paper we delve deep in the Transformer architecture by investigating two of its core components: self-attention and contextual embeddings. In particular, we study the ident…
Attentive Multi-Task Deep Reinforcement Learning
Timo Bram, Gino Brunner, Oliver Richter +1
Sharing knowledge between tasks is vital for efficient learning in a multi-task setting. However, most research so far has focused on the easier case where knowledge transfer is no…
Learning Policies through Quantile Regression
Oliver Richter, Roger Wattenhofer
Policy gradient based reinforcement learning algorithms coupled with neural networks have shown success in learning complex policies in the model free continuous action space contr…
Using State Predictions for Value Regularization in Curiosity Driven Deep Reinforcement Learning
Gino Brunner, Manuel Fritsche, Oliver Richter +1
Learning in sparse reward settings remains a challenge in Reinforcement Learning, which is often addressed by using intrinsic rewards. One promising strategy is inspired by human c…
Teaching a Machine to Read Maps with Deep Reinforcement Learning
Gino Brunner, Oliver Richter, Yuyi Wang +1
The ability to use a 2D map to navigate a complex 3D environment is quite remarkable, and even difficult for many humans. Localization and navigation is also an important problem i…