activity
20242026
collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

Feed-Forward Steering in Transformer Residual Dynamics

Timur Mudarisov, Mikhail Burtsev, Radu State

Attention-only dynamical theories model Transformer residual directions as particles aggregating on a sphere. We extend this framework by incorporating the feed-forward network (FF…

cs.LG2026

Geometry-Guided Layerwise FFN Width Allocation in Transformers

Timur Mudarisov, Mikhail Burtsev, Radu State

Feed-forward networks (FFNs) account for a large fraction of Transformer parameters, yet their hidden width is usually constant across depth. We ask whether this capacity can inste…

cs.LG2026

QeHDC: Hyperdimensional Computing based on Quantum-enhanced binding and SuperClass Construction

Yangjie Xu, Hui Huang, Li Ning +1

Hyperdimensional Computing (HDC) is a robust computational framework inspired by human cognition characterized by simple and efficient operations within high-dimensional vector spa…

cs.LG2026

Limitations of Normalization in Attention Mechanism

Timur Mudarisov, Mikhail Burtsev, Tatiana Petrova +1

This paper investigates the limitations of the normalization in attention mechanisms. We begin with a theoretical framework that enables the identification of the model's selective…

cs.LG2026

Friction-Augmented Drifting Models for Resource-Efficient Domain Translation

Arkadii Kazanskii, Tatiana Petrova, Andrey Ustyuzhanin +3

Single-step generators promise high-fidelity synthesis at a fraction of the inference and training cost of ordinary differential equation (ODE)-based flow models, a central concern…