Dispatcher: A Message-Passing Approach To Language Modelling
arXiv:2105.03994
Abstract
This paper proposes a message-passing mechanism to address language modelling. A new layer type is introduced that aims to substitute self-attention for unidirectional sequence generation tasks. The system is shown to be competitive with existing methods: Given N tokens, the computational complexity is O(N logN) and the memory complexity is O(N) under reasonable assumptions. In the end, the Dispatcher layer is seen to achieve comparable perplexity to prior results while being more efficient.
References in corpus (10)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Convolutional Sequence to Sequence Learning
- Linformer: Self-Attention with Linear Complexity
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Pointer Sentinel Mixture Models
- Big Bird: Transformers for Longer Sequences
- ConvTransformer: A Convolutional Transformer Network for Video Frame Synthesis
- Language Models with Transformers
- An Attention Free Transformer