Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Stabilized Best-of- Training for Neural Combinatorial Optimization
Melveena Jolly, Midhun Xavier
Leader Reward modifies POMO training to emphasize the best trajectory produced by repeated inference. We test a narrow extension: replace its binary leader/non-leader distinction w…
cs.LG2026
Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of- Objective
Melveena Jolly, Midhun Xavier
We study the coupled objective J_K^WOR = E_{S ~ PL-WOR_K}[max_{i in S} R_i]: the expected maximum reward of a size-K Plackett-Luce draw without replacement, the law of Gumbel-Top-K…