Contextual Procurement Auctions with Bandit Learning
arXiv:2607.05813
Abstract
We study repeated procurement auctions in which producers have private costs and the platform must learn the context-dependent value of selecting each producer. We evaluate performance by welfare regret: the cumulative loss in total surplus relative to the full-information efficient rule that knows the context-dependent values and true producer costs. The natural UCB allocation rule achieves welfare regret under truthful bids, but its adaptive, bid-dependent learning path does not by itself ensure truthfulness. To obtain exact incentives, we first design a bid-independent explore-then-commit mechanism with empirical threshold payments; it is dominant-strategy truthful and has regret. We then introduce frozen-payment UCB, which estimates payments from initial bid-independent exploration but continues allocation learning by UCB. Under a truthful-path margin condition, the frozen-payment UCB is approximately truthful with an average per-round deviation gain for fixed , . Under truthful bidding, it achieves welfare regret, matching the UCB rate. A lower bound shows that this regret-incentive tradeoff is tight within the frozen critical-payment class.
15 pages, 1 figure