paper

Data Predictability Shapes Weibull Weight-Scale Growth in Transformer Training

arXiv:2608.23573

Abstract

A trained transformer's weight magnitudes can be summarized by a two-parameter Weibull distribution whose shape is stable across layers and models, so the scale carries most training-induced movement. What corpus property sets how much grows? Using the bigram conditional entropy , a training-free statistic computed before training, we find across controlled corruption families a learning-rate-conditioned law, , where is a matched-budget shuffle baseline. The convex exponent is inherited from an independently measured data-side saturation relation rather than fitted directly to the growth curve. After removing the two per- coefficients, 23 runs spanning an order of magnitude in learning rate collapse onto with unit slope (; direct per- fits are weaker, ). Because is computed before training, the law is a forward predictor: an end-to-end self-validation recovers held-out within-family weight growth with 5.7% relative error. The readout holds at model and per-layer resolutions and across two tested architectures, with the functional form preserved and only the coefficients changing. It also marks its boundary: cross-corpus prediction over-predicts code, implicating redundancy as a second axis of a broader data-to-weight framework.

27 pages, 14 figures, 5 tables. Code and data: https://github.com/tiexinding/NPM-Weibull-public

Data Predictability Shapes Weibull Weight-Scale Growth in Transformer Training · wovepaper