4 papers
Feed-Forward Steering in Transformer Residual Dynamics
Timur Mudarisov, Mikhail Burtsev, Radu State
Attention-only dynamical theories model Transformer residual directions as particles aggregating on a sphere. We extend this framework by incorporating the feed-forward network (FF…
Geometry-Guided Layerwise FFN Width Allocation in Transformers
Timur Mudarisov, Mikhail Burtsev, Radu State
Feed-forward networks (FFNs) account for a large fraction of Transformer parameters, yet their hidden width is usually constant across depth. We ask whether this capacity can inste…
Limitations of Normalization in Attention Mechanism
Timur Mudarisov, Mikhail Burtsev, Tatiana Petrova +1
This paper investigates the limitations of the normalization in attention mechanisms. We begin with a theoretical framework that enables the identification of the model's selective…
Geometric Analysis of Token Selection in Multi-Head Attention
Timur Mudarisov, Mikhal Burtsev, Tatiana Petrova +1
We present a geometric framework for analysing multi-head attention in large language models (LLMs). Without altering the mechanism, we view standard attention through a top-N sele…