Learnability Window in Gated Recurrent Neural Networks
arXiv:2512.05790 · doi:10.1103/843n-yshj
The paper develops a statistical theory describing how the gating mechanisms and optimizers in recurrent neural networks determine the maximum time span over which gradient-based learning can recover temporal dependencies, especially under heavy‑tailed noise.
Abstract
We develop a statistical theory of temporal learnability in recurrent neural networks, quantifying the maximal temporal horizon over which gradient-based learning can recover lag-dependent structure at finite sample size . The theory is built on the effective learning rate envelope , a function that captures how gating mechanisms and adaptive optimizers jointly shape the coupling between state-space dynamics and parameter updates during Backpropagation Through Time. Under heavy-tailed (-stable) fluctuations, where empirical averages concentrate at rate with , the interplay between envelope decay and statistical concentration yields explicit scaling laws for the growth of : logarithmic, polynomial, and exponential temporal learning regimes emerge according to the decay law of . These results identify envelope decay as the key determinant of temporal learnability. Slower attenuation of enlarges , while heavy-tailed fluctuations compress it by weakening statistical concentration. Moreover, envelope geometry outweighs dataset size: slowing the envelope's decay enlarges more than adding data, so more complex architectures that realize slower-decaying envelopes can be more data-efficient than simpler ones. Experiments across multiple gated architectures and optimizers corroborate these structural predictions.
Accepted at Physical Review E