1 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Tamir David Hay, Lior Wolf
In the pursuit of reducing the number of trainable parameters in deep transformer networks, we employ Reinforcement Learning to dynamically select layers during training and tie th…