8 citations · 14 across the 2 of their papers we have counts for
2 papers
cs.LG2017★ 6 cited
Diagonal Rescaling For Neural Networks
Jean Lafond, Nicolas Vasilache, Léon Bottou
We define a second-order neural network stochastic gradient training algorithm whose block-diagonal structure effectively amounts to normalizing the unit activations. Investigating…
cs.CL2017★ 8 cited
Training Language Models Using Target-Propagation
Sam Wiseman, Sumit Chopra, Marc'Aurelio Ranzato +4
While Truncated Back-Propagation through Time (BPTT) is the most popular approach to training Recurrent Neural Networks (RNNs), it suffers from being inherently sequential (making…