1 paper · 1 filter
Alankar Atreya, Devesh Batra, Yoages Kumar Mantri +3
Continual pre-training of large language models must acquire new information without erasing old knowledge. Existing replay methods often choose a global old/new mixture and sample…