Choose a Transformer: Fourier or Galerkin
arXiv:2105.14995
Abstract
In this paper, we apply the self-attention from the state-of-the-art Transformer in Attention Is All You Need for the first time to a data-driven operator learning problem related to partial differential equations. An effort is put together to explain the heuristics of, and to improve the efficacy of the attention mechanism. By employing the operator approximation theory in Hilbert spaces, it is demonstrated for the first time that the softmax normalization in the scaled dot-product attention is sufficient but not necessary. Without softmax, the approximation capacity of a linearized Transformer variant can be proved to be comparable to a Petrov-Galerkin projection layer-wise, and the estimate is independent with respect to the sequence length. A new layer normalization scheme mimicking the Petrov-Galerkin projection is proposed to allow a scaling to propagate through attention layers, which helps the model achieve remarkable accuracy in operator learning tasks with unnormalized data. Finally, we present three operator learning experiments, including the viscid Burgers' equation, an interface Darcy flow, and an inverse interface coefficient identification problem. The newly proposed simple attention-based operator learner, Galerkin Transformer, shows significant improvements in both training cost and evaluation accuracy over its softmax-normalized counterparts.
35 pages, 13 figures. Published as a conference paper at NeurIPS 2021
References in corpus (23)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Sequence to Sequence Learning with Neural Networks
- MLP-Mixer: An all-MLP Architecture for Vision
- Linformer: Self-Attention with Linear Complexity
- Searching for Activation Functions
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- SE(3)-Transformers: 3D Roto-Translation Equivariant Attention Networks
- Understanding and mitigating gradient pathologies in physics-informed neural networks
- Multipole Graph Neural Operator for Parametric Partial Differential Equations
- The Random Feature Model for Input-Output Maps between Banach Spaces
- Random Feature Attention
- On Layer Normalization in the Transformer Architecture
- Pretrained Transformers as Universal Computation Engines
- Multiwavelet-based Operator Learning for Differential Equations
- Are Transformers universal approximators of sequence-to-sequence functions?
- Equivariant Transformer Networks
- Linear Transformers Are Secretly Fast Weight Programmers
- A Cheap Linear Attention Mechanism with Fast Lookups and Fixed-Size Representations
- The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of Transformers
- FMMformer: Efficient and Flexible Transformer via Decomposed Near-field and Far-field Attention
- Transformers are Deep Infinite-Dimensional Non-Mercer Binary Kernel Machines
- Rethinking Neural Operations for Diverse Tasks
- Implicit Kernel Attention
Cited by in corpus (12)
- CenterCLIP: Token Clustering for Efficient Text-Video Retrieval
- Adaptive Fourier Neural Operators: Efficient Token Mixers for Transformers
- Fast Dynamic 1D Simulation of Divertor Plasmas with Neural PDE Surrogates
- Variational operator learning: A unified paradigm marrying training neural operators and solving partial differential equations
- Meta-Auto-Decoder for Solving Parametric Partial Differential Equations
- Multi-scale Time-stepping of Partial Differential Equations with Transformers
- Deciphering and integrating invariants for neural operator learning with various physical mechanisms
- Enhancing generalization in endwall film cooling prediction: Incorporating the superposition principle into transformer-based neural operators
- Predicting Change, Not States: An Alternate Framework for Neural PDE Surrogates
- MCMC-Net: Accelerating Markov Chain Monte Carlo with Neural Networks for Inverse Problems
- Boundary-to-Solution Mapping for Groundwater Flows in a Toth Basin
- Spectral Transform Forms Scalable Transformer