3 papers
stat.ML2026
A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention
Tomohiro Hayase, Ryo Karakida
Length-dependent logit rescaling is widely used to stabilize long-context self-attention, but existing analyses and methods suggest conflicting inverse-temperature laws for the con…
stat.ML2026
Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix
Tomohiro Hayase, Benoît Collins, Ryo Karakida
Self-attention layers have become fundamental building blocks of modern deep neural networks, yet their theoretical understanding remains limited, particularly from the perspective…
cs.LG2026
Free Random Projection for In-Context Reinforcement Learning
Tomohiro Hayase, Benoît Collins, Nakamasa Inoue
Hierarchical inductive biases are hypothesized to promote generalizable policies in reinforcement learning, as demonstrated by explicit hyperbolic latent representations and archit…