3 papers
cs.LG2026
ASAP: Amortized Doubly-Stochastic Attention via Sliced Dual Projection
Huy Tran, Max Milkert, David Hyde
Doubly-stochastic attention has emerged as a transport-based alternative to row-softmax attention, with recent Transformer variants using it to reduce attention sinks and rank coll…
cs.LG2023
Compelling ReLU Networks to Exhibit Exponentially Many Linear Regions at Initialization and During Training
Max Milkert, David Hyde, Forrest Laine
In a neural network with ReLU activations, the number of piecewise linear regions in the output can grow exponentially with depth. However, this is highly unlikely to happen when t…
math.NT2022
Prime Holdout Problems
Max Milkert, Alex Ruchti, Josiah Yoder
This paper introduces prime holdout problems, a problem class related to the Collatz conjecture. After applying a linear function, instead of removing a finite set of prime factors…