5 papers
FourierQK: Spectral Preprocessing of Query-Key Projections Improves Transformer Attention
Athanasios Zeris
FFT-based spectral preprocessing of learned query-key (Q/K) projections substantially improves transformer attention on character-level language modelling. On TinyShakespeare: a fi…
Multiscale POD of Transformer Attention Fields: Scale-Selective Analysis via Morlet Scalogram
Athanasios Zeris
We introduce scale-selective Proper Orthogonal Decomposition (POD) for transformer attention fields, inspired by the use of POD for extracting energetically dominant modes from tur…
Beyond Sinusoids: A Morlet Wavelet Framework for Transformer Positional Encoding
Athanasios Zeris
Standard positional encodings for transformers - sinusoidal and rotary (RoPE) - treat every position as equally local: they encode where a token is, but not how far its positional…
Energy-Gated Attention and Wavelet Positional Encoding: Complementary Inductive Biases for Transformer Attention
Athanasios Zeris
Standard transformer attention computes pairwise token similarity but treats all tokens as equally salient and all positions as equally local, regardless of the informational struc…
Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention
Athanasios Zeris
Standard transformer attention computes pairwise similarity between queries and keys, treating all tokens as equally salient regardless of their intrinsic informational content. In…