4 papers · 1 filter
An Overview of Low-Rank Structures in the Training and Adaptation of Large Models
Laura Balzano, Tianjiao Ding, Benjamin D. Haeffele +5
The substantial computational demands of modern large-scale deep learning present significant challenges for efficient training and deployment. Recent research has revealed a wides…
A Convex Relaxation Approach to Generalization Analysis for Parallel Positively Homogeneous Networks
Uday Kiran Reddy Tadipatri, Benjamin D. Haeffele, Joshua Agterberg +1
We propose a general framework for deriving generalization bounds for parallel positively homogeneous neural networks--a class of neural networks whose input-output map decomposes…
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
Ziyang Wu, Tianjiao Ding, Yifu Lu +6
The attention operator is arguably the key distinguishing factor of transformer architectures, which have demonstrated state-of-the-art performance on a variety of tasks. However,…
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
Yaodong Yu, Sam Buchanan, Druv Pai +7
In this paper, we contend that a natural objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a low-dimension…