4 papers · 1 filter
ASAP: Amortized Doubly-Stochastic Attention via Sliced Dual Projection
Huy Tran, Max Milkert, David Hyde
Doubly-stochastic attention has emerged as a transport-based alternative to row-softmax attention, with recent Transformer variants using it to reduce attention sinks and rank coll…
Efficient Sliced Wasserstein Distance Computation via Adaptive Bayesian Optimization
Manish Acharya, David Hyde
The sliced Wasserstein distance (SW) reduces optimal transport on to a sum of one-dimensional projections, and thanks to this efficiency, it is widely used in geomet…
Compelling ReLU Networks to Exhibit Exponentially Many Linear Regions at Initialization and During Training
Max Milkert, David Hyde, Forrest Laine
In a neural network with ReLU activations, the number of piecewise linear regions in the output can grow exponentially with depth. However, this is highly unlikely to happen when t…
Building Machine Learning Challenges for Anomaly Detection in Science
Elizabeth G. Campolongo, Yuan-Tang Chou, Ekaterina Govorkova +148
Scientific discoveries are often made by finding a pattern or object that was not predicted by the known rules of science. Oftentimes, these anomalous events or objects that do not…