5 papers
Self-Attention at Constant Cost per Token via Symmetry-Aware Taylor Approximation
Franz A. Heinsen, Leo Kozachkov
The most widely used artificial intelligence (AI) models today are Transformers employing self-attention. In its standard form, self-attention incurs costs that increase with conte…
Generalized Orders of Magnitude for Scalable, Parallel, High-Dynamic-Range Computation
Franz A. Heinsen, Leo Kozachkov
Many domains, from deep learning to finance, require compounding real numbers over long sequences, often leading to catastrophic numerical underflow or overflow. We introduce gener…
Parallelizing MCMC Across the Sequence Length
David M. Zoltowski, Skyler Wu, Xavier Gonzalez +2
Markov chain Monte Carlo (MCMC) methods are foundational algorithms for Bayesian inference and probabilistic modeling. However, most MCMC algorithms are inherently sequential and t…
Intrinsic Goals for Autonomous Agents: Model-Based Exploration in Virtual Zebrafish Predicts Ethological Behavior and Whole-Brain Dynamics
Reece Keller, Alyn Kirsch, Felix Pei +3
Autonomy is a hallmark of animal intelligence, enabling adaptive and intelligent behavior in complex environments without relying on external reward or task structure. Existing rei…
Is All Learning (Natural) Gradient Descent?
Lucas Shoji, Kenta Suzuki, Leo Kozachkov
This paper shows that a wide class of effective learning rules -- those that improve a scalar performance measure over a given time window -- can be rewritten as natural gradient d…