4 citations · 4 across the 2 of their papers we have counts for
1 paper · 1 filter
Yao Fu, Rameswar Panda, Xinyao Niu +4
We study the continual pretraining recipe for scaling language models' context lengths to 128K, with a focus on data engineering. We hypothesize that long context modeling, in part…