Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View
arXiv:1906.02762
Abstract
The Transformer architecture is widely used in natural language processing. Despite its success, the design principle of the Transformer remains elusive. In this paper, we provide a novel perspective towards understanding the architecture: we show that the Transformer can be mathematically interpreted as a numerical Ordinary Differential Equation (ODE) solver for a convection-diffusion equation in a multi-particle dynamic system. In particular, how words in a sentence are abstracted into contexts by passing through the layers of the Transformer can be interpreted as approximating multiple particles' movement in the space using the Lie-Trotter splitting scheme and the Euler's method. Given this ODE's perspective, the rich literature of numerical analysis can be brought to guide us in designing effective structures beyond the Transformer. As an example, we propose to replace the Lie-Trotter splitting scheme by the Strang-Marchuk splitting scheme, a scheme that is more commonly used and with much lower local truncation errors. The Strang-Marchuk splitting scheme suggests that the self-attention and position-wise feed-forward network (FFN) sub-layers should not be treated equally. Instead, in each layer, two position-wise FFN sub-layers should be used, and the self-attention sub-layer is placed in between. This leads to a brand new architecture. Such an FFN-attention-FFN layer is "Macaron-like", and thus we call the network with this new architecture the Macaron Net. Through extensive experiments, we show that the Macaron Net is superior to the Transformer on both supervised and unsupervised learning tasks. The reproducible codes and pretrained models can be found at https://github.com/zhuohan123/macaron-net
References in corpus (9)
- Pay Less Attention with Lightweight and Dynamic Convolutions
- Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
- Bridging the Gaps Between Residual Learning, Recurrent Neural Networks and Visual Cortex
- Weighted Transformer Network for Machine Translation
- Fixup Initialization: Residual Learning Without Normalization
- Understanding Back-Translation at Scale
- Nonlocal Neural Networks, Nonlocal Diffusion and Nonlocal Modeling
- Dynamically Unfolding Recurrent Restorer: A Moving Endpoint Control Method for Image Restoration
- Transport Analysis of Infinitely Deep Neural Network
Cited by in corpus (26)
- Conformer: Convolution-augmented Transformer for Speech Recognition
- A Review on Deep Learning in Medical Image Reconstruction
- Is Attention Better Than Matrix Decomposition?
- A Comprehensive Survey of Machine Learning Applied to Radar Signal Processing
- A Mean-field Analysis of Deep ResNet and Beyond: Towards Provable Optimization Via Overparameterization From Depth
- ODE Transformer: An Ordinary Differential Equation-Inspired Model for Neural Machine Translation
- Neural Machine Translation: Challenges, Progress and Future
- Unification of Symmetries Inside Neural Networks: Transformer, Feedforward and Neural ODE
- Does Simultaneous Speech Translation need Simultaneous Models?
- A Mathematical Theory of Attention
- Mask Attention Networks: Rethinking and Strengthen Transformer
- Improved Mask-CTC for Non-Autoregressive End-to-End ASR
- Darts-Conformer: Towards Efficient Gradient-Based Neural Architecture Search For End-to-End ASR
- Non-autoregressive End-to-end Speech Translation with Parallel Autoregressive Rescoring
- Redesigning the Transformer Architecture with Insights from Multi-particle Dynamical Systems
- DDPNOpt: Differential Dynamic Programming Neural Optimizer
- IOT: Instance-wise Layer Reordering for Transformer Structures
- Staircase Attention for Recurrent Processing of Sequences
- Interpolation between Residual and Non-Residual Networks
- DelightfulTTS: The Microsoft Speech Synthesis System for Blizzard Challenge 2021
- On the Regularity of Attention
- Transformers are Deep Infinite-Dimensional Non-Mercer Binary Kernel Machines
- Adversarial Joint Training with Self-Attention Mechanism for Robust End-to-End Speech Recognition
- ResIST: Layer-Wise Decomposition of ResNets for Distributed Training
- UniDrop: A Simple yet Effective Technique to Improve Transformer without Extra Cost
- Implicit regularization of deep residual networks towards neural ODEs