2.7k citations · 2.7k across the 3 of their papers we have counts for
1 paper · 1 filter
Igor Molybog, Peter Albert, Moya Chen +14
We present a theory for the previously unexplained divergent behavior noticed in the training of large language models. We argue that the phenomenon is an artifact of the dominant…