1 paper · 1 filter
Arthur Renard, Franck Gabriel, Valentin Hartmann +1
We present Frost Training, a method for improving Monte Carlo-based policy optimization for a large family of LLM-as-a-judge tasks called Cross-Entropy Games. The key idea is to ex…