The Boltzmann structure of sampling: Intrinsic -value and its emergent closed-form expression
arXiv:2608.14608
Abstract
We consider observables whose realizations in a sampled dataset are restricted, for example by measurement resolution, to a finite set of distinguishable categories within their possibly infinite theoretical domain. Given the probabilities of observable categories, we study the coarse-grained probability mass of families of possible datasets generated through a forward sampling process. The combinatorial construction induces an intrinsically discrete -value defined directly from the sampling process rather than through additional probabilistic structure on observables. Specifying the sample means of arbitrary functions defines a linear family of datasets whose probability mass is obtained as a weighted sum over integer lattice points contained within the associated polyhedron. To overcome the intractable large- combinatorics, we derive via saddle-point techniques a density approximating these probability masses in the continuum limit of forward sampling within the multinomial universality class. As a demonstration, we consider conditional sampling, where structural means are fixed while means vary over admissible datasets. The information geometry emerging from the saddle-point density, together with the spherical symmetry arising at large from the intrinsic -value construction, enables efficient computation of the -value in the Laplace approximation via the distribution. The resulting statistic is given by the semi-analytic expression times the Kullback-Leibler divergence between the information projections associated with the corresponding structural and observed linear families. These projections can be computed efficiently via standard numerical routines converging for sufficiently well-behaved sample means.