Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Initialization of Large Language Models via Reparameterization to Mitigate Loss Spikes
Kosuke Nishida, Kyosuke Nishida, Kuniko Saito
Loss spikes, a phenomenon in which the loss value diverges suddenly, is a fundamental issue in the pre-training of large language models. This paper supposes that the non-uniformit…
cs.CL2021
Task-adaptive Pre-training of Language Models with Word Embedding Regularization
Kosuke Nishida, Kyosuke Nishida, Sen Yoshida
Pre-trained language models (PTLMs) acquire domain-independent linguistic knowledge through pre-training with massive textual resources. Additional pre-training is effective in ada…