Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?
Fırat Öncel, Matthias Bethge, Beyza Ermis +3
In the last decade, the generalization and adaptation abilities of deep learning models were typically evaluated on fixed training and test distributions. Contrary to traditional d…
cs.CL2024
Investigating Continual Pretraining in Large Language Models: Insights and Implications
Çağatay Yıldız, Nishaanth Kanna Ravichandran, Nitin Sharma +2
Continual learning (CL) in large language models (LLMs) is an evolving domain that focuses on developing efficient and sustainable training strategies to adapt models to emerging k…