collaborators

5 papers

cs.LG2026

FourierQK: Spectral Preprocessing of Query-Key Projections Improves Transformer Attention

Athanasios Zeris

FFT-based spectral preprocessing of learned query-key (Q/K) projections substantially improves transformer attention on character-level language modelling. On TinyShakespeare: a fi…

physics.flu-dyn2026

Multiscale POD of Transformer Attention Fields: Scale-Selective Analysis via Morlet Scalogram

Athanasios Zeris

We introduce scale-selective Proper Orthogonal Decomposition (POD) for transformer attention fields, inspired by the use of POD for extracting energetically dominant modes from tur…

cs.LG2026

Beyond Sinusoids: A Morlet Wavelet Framework for Transformer Positional Encoding

Athanasios Zeris

Standard positional encodings for transformers - sinusoidal and rotary (RoPE) - treat every position as equally local: they encode where a token is, but not how far its positional…

cs.LG2026

Energy-Gated Attention and Wavelet Positional Encoding: Complementary Inductive Biases for Transformer Attention

Athanasios Zeris

Standard transformer attention computes pairwise token similarity but treats all tokens as equally salient and all positions as equally local, regardless of the informational struc…

cs.LG2026

Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention

Athanasios Zeris

Standard transformer attention computes pairwise similarity between queries and keys, treating all tokens as equally salient regardless of their intrinsic informational content. In…