1 paper
Jin-woo Lee, Junhwa Choi, Bongkyu Hwang +8
We revisit continual pre-training for large language models and argue that progress now depends more on scaling the right structure than on scaling parameters alone. We introduce S…