Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Attention as In-Context Empirical Bayes: A Two-Stage View via Particle Dynamics
Matthew Smart, Soumya Ganguly, Nilava Metya +2
We study minimal attention-only transformers under all-token corruption and show they admit a two-stage empirical Bayes interpretation. A single attention step computes a kernel-we…
cs.LG2025
In-context denoising with one-layer transformers: connections between attention and associative memory retrieval
Matthew Smart, Alberto Bietti, Anirvan M. Sengupta
We introduce in-context denoising, a task that refines the connection between attention-based architectures and dense associative memory (DAM) networks, also known as modern Hopfie…