1 paper
Guhan Chen, Songtao Tian, Bohan Li +3
Iterative preference optimization is essential for aligning Large Language Models on mathematical reasoning tasks, yet its efficiency is often throttled by signal scarcity: as the…