4 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.CL2026
Self-Improving Pretraining: using post-trained models to pretrain better models
Ellen Xiaoqing Tan, Jack Lanchantin, Shehzaad Dhuliawala +9
Large language models are classically trained in stages: pretraining on raw text followed by post-training for instruction following and reasoning. However, this separation creates…
cs.LG2021★ 4 cited
Characterizing and addressing the issue of oversmoothing in neural autoregressive sequence modeling
Ilia Kulikov, Maksim Eremeev, Kyunghyun Cho
Neural autoregressive sequence models smear the probability among many possible sequences including degenerate ones, such as empty or repetitive sequences. In this work, we tackle…