Buying Data of Unknown Quality: Statistical Information Procurement Auctions
arXiv:2604.08821
Abstract
We study statistical parameter estimation in the setting of data markets when providers differ in both provision costs and estimation-relevant data quality. We define a cost-per-information score that summarizes each provider's provision cost per unit of information about the buyer's estimation objective. When quality is known ex ante, we describe a second-score procurement mechanism that ranks providers by this score, and endogenously chooses both a provider and a sample size while making truthful cost reports optimal. We then turn to the more realistic setting where data quality is private, and can only be assessed noisily via the delivered data. In this setting, we propose a simple mechanism that augments the second-score rule with a lenient ex post statistical test of the reported quality. We prove that, under mild conditions, the mechanism admits near-truthful reports that become approximately optimal and individually rational as the procured sample size grows. Under these near-truthful reports, the buyer asymptotically recovers the performance of the corresponding known-quality second-score mechanism. Our analysis highlights how the choice of verification test and the buyer's accuracy-cost tradeoff jointly shape participation and misreporting incentives in data markets.