55 citations · 55 across the 1 of their papers we have counts for
2 papers
cs.LG2022★ 55 cited
Reducing Activation Recomputation in Large Transformer Models
Vijay Korthikanti, Jared Casper, Sangkug Lym +4
Training large transformer models is one of the most important computational challenges of modern AI. In this paper, we show how to significantly accelerate training of large trans…
cs.NE2018
Sparse Persistent RNNs: Squeezing Large Recurrent Networks On-Chip
Feiwen Zhu, Jeff Pool, Michael Andersch +2
Recurrent Neural Networks (RNNs) are powerful tools for solving sequence-based problems, but their efficacy and execution time are dependent on the size of the network. Following r…