Are queries and keys always relevant? A case study on Transformer wave functions
arXiv:2405.18874 · doi:10.1088/2632-2153/ada1a0
Abstract
The dot product attention mechanism, originally designed for natural language processing tasks, is a cornerstone of modern Transformers. It adeptly captures semantic relationships between word pairs in sentences by computing a similarity overlap between queries and keys. In this work, we explore the suitability of Transformers, focusing on their attention mechanisms, in the specific domain of the parametrization of variational wave functions to approximate ground states of quantum many-body spin Hamiltonians. Specifically, we perform numerical simulations on the two-dimensional - Heisenberg model, a common benchmark in the field of quantum many-body systems on lattice. By comparing the performance of standard attention mechanisms with a simplified version that excludes queries and keys, relying solely on positions, we achieve competitive results while reducing computational cost and parameter usage. Furthermore, through the analysis of the attention maps generated by standard attention mechanisms, we show that the attention weights become effectively input-independent at the end of the optimization. We support the numerical results with analytical calculations, providing physical insights of why queries and keys should be, in principle, omitted from the attention mechanism when studying large systems.
10 pages, 5 figures, 1 table
References in corpus (7)
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- On the origin of long-range correlations in texts
- A simple linear algebra identity to optimize Large-Scale Neural Network Quantum States
- Optimizing Design Choices for Neural Quantum States
- Ab-initio variational wave functions for the time-dependent many-electron Schrödinger equation
- Fine-tuning Neural Network Quantum States
- Vision Transformers provably learn spatial structure
Cited by in corpus (3)
- Transformer Wave Function for two dimensional frustrated magnets: emergence of a Spin-Liquid Phase in the Shastry-Sutherland Model
- Design principles of deep translationally-symmetric neural quantum states for frustrated magnets
- Comparing Symmetrized Determinant Neural Quantum States for the Hubbard Model