paper

Closing the Theory-Practice Gap in Spiking Transformers via Effective Dimension

arXiv:2604.15769

Abstract

Spiking transformers achieve competitive accuracy with conventional transformers while offering - energy efficiency on neuromorphic hardware, yet no theoretical framework guides their design. This paper establishes the first comprehensive expressivity theory for spiking self-attention. We prove that spiking attention with Leaky Integrate-and-Fire neurons is a universal approximator of continuous permutation-equivariant functions, providing explicit spike circuit constructions including a novel lateral inhibition network for softmax normalization with proven convergence. We derive tight spike-count lower bounds via rate-distortion theory: -approximation requires spikes, with rigorous information-theoretic derivation. Our key insight is input-dependent bounds using measured effective dimensions (-- for CIFAR/ImageNet), explaining why timesteps suffice despite worst-case predictions. We provide concrete design rules with calibrated constants (, 95\% CI: ). Experiments on Spikformer, QKFormer, and SpikingResformer across vision and language benchmarks validate predictions with (). Our framework provides the first principled foundation for neuromorphic transformer design.

6 pages, 3 figures, 7 tables

Closing the Theory-Practice Gap in Spiking Transformers via Effective Dimension · wovepaper