1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Shiyue Zhang, Shijie Wu, Ozan Irsoy +4
Autoregressive language models are trained by minimizing the cross-entropy of the model distribution Q relative to the data distribution P -- that is, minimizing the forward cross-…