Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
From Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASP
Michael Rizvi-Martel, Satwik Bhattamishra, Guillaume Rabusseau +1
A theoretical understanding of Transformers is crucial to better understand the capacities and limitations of large language models (LLMs). There is much work analyzing the express…
cs.LG2026
On the Role of Depth in the Expressivity of RNNs
Maude Lizaire, Michael Rizvi-Martel, Éric Dupuis +1
The benefits of depth in feedforward neural networks are well known: composing multiple layers of linear transformations with nonlinear activations enables complex computations. Wh…
cs.LG2024
A Tensor Decomposition Perspective on Second-order RNNs
Maude Lizaire, Michael Rizvi-Martel, Marawan Gamal Abdel Hameed +1
Second-order Recurrent Neural Networks (2RNNs) extend RNNs by leveraging second-order interactions for sequence modelling. These models are provably more expressive than their firs…