paper

On the Statistical Complexity of Sample Amplification

arXiv:2201.04315

Abstract

The ``sample amplification'' problem formalizes the following question: Given i.i.d. samples drawn from an unknown distribution , when is it possible to produce a larger set of samples which cannot be distinguished from i.i.d. samples drawn from ? In this work, we provide a firm statistical foundation for this problem by deriving generally applicable amplification procedures, lower bound techniques and connections to existing statistical notions. Our techniques apply to a large class of distributions including the exponential family, and establish a rigorous connection between sample amplification and distribution learning.

To appear in the Annals of Statistics