activity
20182022
most citedScatterbrain: Unifying Sparse and Low-rank Attention Approximation

9 citations · 14 across the 2 of their papers we have counts for

collaborators

6 papers

cs.LG20225 cited

Interpreting Neural Networks through the Polytope Lens

Sid Black, Lee Sharkey, Leo Grinsztajn +8

Mechanistic interpretability aims to explain what a neural network has learned at a nuts-and-bolts level. What are the fundamental primitives of neural network representations? Pre…

cs.LG20219 cited

Scatterbrain: Unifying Sparse and Low-rank Attention Approximation

Beidi Chen, Tri Dao, Eric Winsor +3

Recent advances in efficient Transformers have exploited either the sparsity or low-rank properties of attention matrices to reduce the computational and memory bottlenecks of mode…

math.CO2019

Generalized Lyndon Factorizations of Infinite Words

Amanda Burcroff, Eric Winsor

A generalized lexicographic order on words is a lexicographic order where the total order of the alphabet depends on the position of the comparison. A generalized Lyndon word is a…

math.NT2019

A Refined Conjecture for the Variance of Gaussian Primes Across Sectors

Ryan C. Chen, Yujin H. Kim, Jared D. Lichtman +6

We derive a refined conjecture for the variance of Gaussian primes across sectors, with a power saving error term, by applying the L-functions Ratios Conjecture. We observe a bifur…

math.NT2018

Limiting Distributions in Generalized Zeckendorf Decompositions

Alexandre Gueganic, Granger Carty, Yujin H. Kim +5

An equivalent definition of the Fibonacci numbers is that they are the unique sequence such that every integer can be written uniquely as a sum of non-adjacent terms. We can view t…

math-ph2018

Spectral Statistics of Non-Hermitian Random Matrix Ensembles

Ryan C. Chen, Yujin H. Kim, Jared D. Lichtman +3

Recently Burkhardt et. al. introduced the -checkerboard random matrix ensembles, which have a split limiting behavior of the eigenvalues (in the limit all but of the eigenva…