2 papers
cs.LG2020
Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout
Zhao Chen, Jiquan Ngiam, Yanping Huang +4
The vast majority of deep models use multiple gradient signals, typically corresponding to a sum of multiple loss terms, to update a shared set of trainable weights. However, these…
cs.LG2018
Learning Longer-term Dependencies in RNNs with Auxiliary Losses
Trieu H. Trinh, Andrew M. Dai, Minh-Thang Luong +1
Despite recent advances in training recurrent neural networks (RNNs), capturing long-term dependencies in sequences remains a fundamental challenge. Most approaches use backpropaga…