Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Reusing Overtrained Language Models Saturates Scaling
Seng Pei Liew, Takuya Kato
Reusing pretrained base models for further pretraining, such as continual pretraining or model growth, is promising at reducing the cost of training language models from scratch. H…
cs.CL2024
Large Vocabulary Size Improves Large Language Models
Sho Takase, Ryokan Ri, Shun Kiyono +1
This paper empirically investigates the relationship between subword vocabulary size and the performance of large language models (LLMs) to provide insights on how to define the vo…