paper

Stabilized Best-of- Training for Neural Combinatorial Optimization

arXiv:2608.00296

Abstract

Leader Reward modifies POMO training to emphasize the best trajectory produced by repeated inference. We test a narrow extension: replace its binary leader/non-leader distinction with a stabilized rank signal indexed by a sampling budget . With the POMO architecture, 3,050-epoch schedule, and TSP-100 test set held fixed, the Leader Reward reimplementation obtains under 100-start, 8-augmentation greedy decoding, matching the reported at its displayed precision. Under independent sampling, the stabilized recipe lowers realized Best-of-8 cost in all three paired training seeds: versus . This observation is estimation-only and decoder-specific: three seeds are below the six-seed testing floor, Leader Reward is better at sampled , and it remains slightly better under its original augmented-greedy protocol. We make no unbiased-estimator, universal superiority, or state-of-the-art claim.

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization · wovepaper