2 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Jin-woo Lee, Junhwa Choi, Bongkyu Hwang +8
We revisit continual pre-training for large language models and argue that progress now depends more on scaling the right structure than on scaling parameters alone. We introduce S…