1 paper
Kunal Samanta, Ari Holtzman, Peter West
The alignment of language models is typically studied through the lens of capability benchmarks, but the dynamics of how models change during post-training remain poorly understood…