paper

One Spectrum, Two Resources: Data-Memory Scaling in Autoregressive Prediction

arXiv:2609.13500

Abstract

How much learned memory is needed to benefit from more data? We show that the two resources are governed by one predictive-energy spectrum in a positive-entropy autoregressive retrieval source. Each coordinate contributes its query probability times the squared radius of its unknown logit. Writing for the resulting energy spectrum, we prove the minimax law for prediction blocks and a learned state with at most values. Data set the resolution ; memory sets the level reached by optimal bit allocation. The complete curve also recovers the positive spectrum. Energy-dimension pairing is essential: two causal sources with identical block-energy and block-dimension marginals have different data and memory exponents. A masked query-key attention head learns the route and values, realizing the law with explicit routing, format, and arithmetic errors. Further results give exponent-adaptive allocation, finite-precision realization, and compute-precision laws under two-sided arithmetic assumptions. Experiments recover the data-memory collapse and coupling exponents, explain the routing and allocation mechanisms, and examine weight-only quantization across six pretrained-model scales.

66 pages, 11 figures, 8 tables

One Spectrum, Two Resources: Data-Memory Scaling in Autoregressive Prediction · wovepaper