2 papers
cs.LG2025
Compact Recurrent Transformer with Persistent Memory
Edison Mucllari, Zachary Daniels, David Zhang +1
The Transformer architecture has shown significant success in many language processing and visual tasks. However, the method faces challenges in efficiently scaling to long sequenc…
cs.LG2024
Preconditioning for Accelerated Gradient Descent Optimization and Regularization
Qiang Ye
Accelerated training algorithms, such as adaptive learning rates (or preconditioning) and various normalization methods, are widely used but not fully understood. When regularizati…