4 papers
Data-Aware Random Feature Kernel for Transformers
Amirhossein Farzam, Hossein Mobahi, Nolan Andrew Miller +1
Transformers excel across domains, yet their quadratic attention complexity poses a barrier to scaling. Random-feature attention, as in Performers, can reduce this cost to linear i…
Time Matters: Scaling Laws for Any Budget
Itay Inbar, Luke Sernau
A primary cost driver for training large models is wall-clock training time. We show that popular time estimates based on FLOPs are poor estimates, and construct a more accurate pr…
All Random Features Representations are Equivalent
Luke Sernau, Silvano Bonacina, Rif A. Saurous
Random features are a powerful technique for rewriting positive-definite kernels as linear products. They bring linear tools to bear in important nonlinear domains like KNNs and at…
Infinite Width Models That Work: Why Feature Learning Doesn't Matter as Much as You Think
Luke Sernau
Common infinite-width architectures such as Neural Tangent Kernels (NTKs) have historically shown weak performance compared to finite models. This is usually attributed to the abse…