171 citations · 244 across the 10 of their papers we have counts for
1 paper · 1 filter
Daniel Kang, Tatsunori Hashimoto
Neural language models are usually trained to match the distributional properties of a large-scale corpus by minimizing the log loss. While straightforward to optimize, this approa…