activity
20202022
most citedScatterbrain: Unifying Sparse and Low-rank Attention Approximation

9 citations · 18 across the 4 of their papers we have counts for

collaborators

5 papers

cs.LG20224 cited

Bypass Exponential Time Preprocessing: Fast Neural Network Training via Weight-Data Correlation Preprocessing

Josh Alman, Jiehao Liang, Zhao Song +2

Over the last decade, deep neural networks have transformed our society, and they are already widely applied in various machine learning applications. State-of-art deep neural netw…

cs.LG20219 cited

Scatterbrain: Unifying Sparse and Low-rank Attention Approximation

Beidi Chen, Tri Dao, Eric Winsor +3

Recent advances in efficient Transformers have exploited either the sparsity or low-rank properties of attention matrices to reduce the computational and memory bottlenecks of mode…

cs.LG20213 cited

Does Preprocessing Help Training Over-parameterized Neural Networks?

Zhao Song, Shuo Yang, Ruizhe Zhang

Deep neural networks have achieved impressive performance in many areas. Designing a fast and provable method for training neural networks is a fundamental question in machine lear…

cs.DS20212 cited

Fast Sketching of Polynomial Kernels of Polynomial Degree

Zhao Song, David P. Woodruff, Zheng Yu +1

Kernel methods are fundamental in machine learning, and faster algorithms for kernel approximation provide direct speedups for many core tasks in machine learning. The polynomial k…

cs.CG2020

Metric Transforms and Low Rank Matrices via Representation Theory of the Real Hyperrectangle

Josh Alman, Timothy Chu, Gary Miller +3

In this paper, we develop a new technique which we call representation theory of the real hyperrectangle, which describes how to compute the eigenvectors and eigenvalues of certain…