1 paper · 1 filter
Kunal Samanta, Ari Holtzman, Peter West
The alignment of language models is typically studied through the lens of capability benchmarks, but the dynamics of how models change during post-training remain poorly understood…