6 papers
InfoFlow: A Framework for Multi-Layer Transformer Analysis
Penghao Yu, Haotian Jiang, Zeyu Bao +1
While the approximation properties of single-layer Transformer architectures have been studied in recent works, a rigorous theoretical understanding of the multi-layer setting rema…
The Effect of Attention Head Count on Transformer Approximation
Penghao Yu, Haotian Jiang, Zeyu Bao +2
Transformer has become the dominant architecture for sequence modeling, yet a detailed understanding of how its structural parameters influence expressive power remains limited. In…
Allocation of Parameters in Transformers
Ruoxi Yu, Haotian Jiang, Jingpu Cheng +3
Transformers have achieved remarkable successes across a wide range of applications, yet the theoretical foundation of their model efficiency remains underexplored. In this work, w…
The Effect of Depth on the Expressivity of Deep Linear State-Space Models
Zeyu Bao, Penghao Yu, Haotian Jiang +1
Deep state-space models (SSMs) have gained increasing popularity in sequence modelling. While there are numerous theoretical investigations of shallow SSMs, how the depth of the SS…
Numerical Investigation of Sequence Modeling Theory using Controllable Memory Functions
Haotian Jiang, Zeyu Bao, Shida Wang +1
The evolution of sequence modeling architectures, from recurrent neural networks and convolutional models to Transformers and structured state-space models, reflects ongoing effort…
Approximation Rate of the Transformer Architecture for Sequence Modeling
Haotian Jiang, Qianxiao Li
The Transformer architecture is widely applied in sequence modeling applications, yet the theoretical understanding of its working principles remains limited. In this work, we inve…