3 papers
cs.CL2026
Reusing Overtrained Language Models Saturates Scaling
Seng Pei Liew, Takuya Kato
Reusing pretrained base models for further pretraining, such as continual pretraining or model growth, is promising at reducing the cost of training language models from scratch. H…
cs.LG2025
Scaling Laws for Upcycling Mixture-of-Experts Language Models
Seng Pei Liew, Takuya Kato, Sho Takase
Pretraining large language models (LLMs) is resource-intensive, often requiring months of training time even with high-end GPU clusters. There are two approaches of mitigating such…
cs.CL2025
Large Vocabulary Size Improves Large Language Models
Sho Takase, Ryokan Ri, Shun Kiyono +1
This paper empirically investigates the relationship between subword vocabulary size and the performance of large language models (LLMs) to provide insights on how to define the vo…