activity
20242026
collaborators

5 papers

cs.LG2026

Self-Attention at Constant Cost per Token via Symmetry-Aware Taylor Approximation

Franz A. Heinsen, Leo Kozachkov

The most widely used artificial intelligence (AI) models today are Transformers employing self-attention. In its standard form, self-attention incurs costs that increase with conte…

cs.LG2025

Generalized Orders of Magnitude for Scalable, Parallel, High-Dynamic-Range Computation

Franz A. Heinsen, Leo Kozachkov

Many domains, from deep learning to finance, require compounding real numbers over long sequences, often leading to catastrophic numerical underflow or overflow. We introduce gener…

stat.CO2025

Parallelizing MCMC Across the Sequence Length

David M. Zoltowski, Skyler Wu, Xavier Gonzalez +2

Markov chain Monte Carlo (MCMC) methods are foundational algorithms for Bayesian inference and probabilistic modeling. However, most MCMC algorithms are inherently sequential and t…

q-bio.NC2025

Intrinsic Goals for Autonomous Agents: Model-Based Exploration in Virtual Zebrafish Predicts Ethological Behavior and Whole-Brain Dynamics

Reece Keller, Alyn Kirsch, Felix Pei +3

Autonomy is a hallmark of animal intelligence, enabling adaptive and intelligent behavior in complex environments without relying on external reward or task structure. Existing rei…

cs.LG2024

Is All Learning (Natural) Gradient Descent?

Lucas Shoji, Kenta Suzuki, Leo Kozachkov

This paper shows that a wide class of effective learning rules -- those that improve a scalar performance measure over a given time window -- can be rewritten as natural gradient d…