9 papers
Contribution Weights: A Geometrical Analysis of Self-Attention Transformers
Harry Jake Cunningham, Nicola Muca Cirone
Analyzing attention weights has become a standard approach for interpreting the information flow of Large Language Models (LLMs). However, this approach has significant limitations…
Structured Linear CDEs: Maximally Expressive and Parallel-in-Time Sequence Models
Benjamin Walker, Lingyi Yang, Nicola Muca Cirone +2
This work introduces Structured Linear Controlled Differential Equations (SLiCEs), a unifying framework for sequence models with structured, input-dependent state-transition matric…
Fixed-Point RNNs: Interpolating from Diagonal to Dense
Sajad Movahedi, Felix Sarnthein, Nicola Muca Cirone +1
Linear recurrent neural networks (RNNs) and state-space models (SSMs) such as Mamba have become promising alternatives to softmax-attention as sequence mixing layers in Transformer…
Kernel Learning for Mean-Variance Trading Strategies
Owen Futter, Nicola Muca Cirone, Blanka Horvath
In this article, we develop a kernel-based framework for constructing dynamic, pathdependent trading strategies under a mean-variance optimisation criterion. Building on the theore…
Genus expansion for non-linear random matrix ensembles with applications to neural networks
Nicola Muca Cirone, Jad Hamdan, Cristopher Salvi
We present a unified approach to studying certain non-linear random matrix ensembles and associated random neural networks at initialization. This begins with a novel series expans…
ParallelFlow: Parallelizing Linear Transformers via Flow Discretization
Nicola Muca Cirone, Cristopher Salvi
We present a theoretical framework for analyzing linear attention models through matrix-valued state space models (SSMs). Our approach, Parallel Flows, provides a perspective that…