2 papers
cs.LG2026
An Overview of Low-Rank Structures in the Training and Adaptation of Large Models
Laura Balzano, Tianjiao Ding, Benjamin D. Haeffele +5
The substantial computational demands of modern large-scale deep learning present significant challenges for efficient training and deployment. Recent research has revealed a wides…
cs.LG2024
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
Ziyang Wu, Tianjiao Ding, Yifu Lu +6
The attention operator is arguably the key distinguishing factor of transformer architectures, which have demonstrated state-of-the-art performance on a variety of tasks. However,…