4 papers
Direct Quantized Training of Language Models with Stochastic Rounding
Kaiyan Zhao, Tsuguchika Tabaru, Kenichi Kobayashi +3
Although recent quantized Large Language Models (LLMs), such as BitNet, have paved the way for significant reduction in memory usage during deployment with binary or ternary weight…
Velocity dependence of the mass modifications of and mesons in 12 GeV reactions
Wataru Nakai, Kazuya Aoki, Junsei Chiba +26
This study measured the invariant mass spectra of and mesons in the decay channel for 12 GeV (12.9 GeV/) and reactions ($\sqrt{…
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
Istabrak Abbes, Gopeshh Subbaraj, Matthew Riemer +6
Training large language models (LLMs) typically involves pre-training on massive corpora, only to restart the process entirely when new data becomes available. A more efficient and…
Analysis of spectral modification of mesons at finite density using a transport approach in the 12 GeV pA reactions
PS E325 Collaboration, Masaya Ichikawa, Philipp Gubler +26
The hadron spectrum at finite density is an important observable for exploring the origin of hadron masses. In the KEK-PS E325 experiment, the di-electron decays of phi mesons insi…