3 papers
cs.LG2026
Gated MLPs as Symmetry-Broken Rank-1 Bilinear Attention
Nathan Breslow
We show that the conventional gated MLP can be viewed as a rank-1 approximation to a bilinear attention mechanism with two distinct factors corresponding to the query and the key.…
cs.LG2026
Genomic Next-Token Predictors are In-Context Learners
Nathan Breslow, Aayush Mishra, Mahler Revsine +3
In-context learning (ICL) -- the capacity of a model to infer and apply abstract patterns from examples provided within its input -- has been extensively studied in large language…
cs.LG2025
Channel-Wise MLPs Improve the Generalization of Recurrent Convolutional Networks
Nathan Breslow
We investigate the impact of channel-wise mixing via multi-layer perceptrons (MLPs) on the generalization capabilities of recurrent convolutional networks. Specifically, we compare…