2 papers
cs.LG2026
Unsupervised Process Reward Models
Artyom Gadetsky, Maxim Kodryan, Siba Smarak Panigrahi +2
Process Reward Models (PRMs) are a powerful mechanism for steering large language model reasoning by providing fine-grained, step-level supervision. However, this effectiveness com…
cs.LG2024
Where Do Large Learning Rates Lead Us?
Ildus Sadrtdinov, Maxim Kodryan, Eduard Pokonechny +2
It is generally accepted that starting neural networks training with large learning rates (LRs) improves generalization. Following a line of research devoted to understanding this…