3 citations · 3 across the 1 of their papers we have counts for
3 papers
Memory in humans and deep language models: Linking hypotheses for model augmentation
Omri Raccah, Phoebe Chen, Ted L. Willke +2
The computational complexity of the self-attention mechanism in Transformer models significantly limits their ability to generalize over long temporal durations. Memory-augmentatio…
Slower is Better: Revisiting the Forgetting Mechanism in LSTM for Slower Information Decay
Hsiang-Yun Sherry Chien, Javier S. Turek, Nicole Beckage +3
Sequential information contains short- to long-range dependencies; however, learning long-timescale information has been a challenge for recurrent neural networks. Despite improvem…
Multi-timescale Representation Learning in LSTM Language Models
Shivangi Mahto, Vy A. Vo, Javier S. Turek +1
Language models must capture statistical dependencies between words at timescales ranging from very short to very long. Earlier work has demonstrated that dependencies in natural l…