Publications (18)
Factored Gossip DiLoCo: Reducing Blocking Communication in DiLoCo
Chamin Hewa Koneputugodage, Thalaiyasingam Ajanthan, Sameera Ramasinghe +7
To make large-scale distributed training practical outside high-bandwidth datacenters, we must reduce blocking, high-volume synchronization. While DiLoCo communicates infrequently,…
Taming Curvature: Architecture Warm-Up for Stable Transformer Training
Sameera Ramasinghe, Ajanthan Thalaiyasingam, Hadi Mohaghegh Dolatabadi +6
Training billion-parameter Transformers is often brittle, with transient loss spikes and divergence that waste compute. Even though the recently developed Edge of Stability (EoS) t…
Effective field theory analysis of the Coulomb breakup of the one-neutron halo nucleus 19C
Pierre Capel, Daniel R. Phillips, Andrew Andis +29
We analyse the Coulomb breakup of 19C measured at 67A MeV at RIKEN. We use the Coulomb-Corrected Eikonal (CCE) approximation to model the reaction and describe the one-neutron halo…
Low energy neutron background in deep underground laboratories
Andreas Best, Joachim Gorres, Matthias Junker +6
The natural neutron background influences the maximum achievable sensitivity in most deep underground nuclear, astroparticle and double-beta decay physics experiments. Reliable neu…
NuMuon: Nuclear-Norm-Constrained Muon for Compressible LLM Training
Hadi Mohaghegh Dolatabadi, Thalaiyasingam Ajanthan, Sameera Ramasinghe +7
The rapid progress of large language models (LLMs) is increasingly constrained by memory and deployment costs, motivating compression methods for practical deployment. Many state-o…
Unextractable Protocol Models: Collaborative Training and Inference without Weight Materialization
Alexander Long, Chamin Hewa Koneputugodage, Thalaiyasingam Ajanthan +5
We consider a decentralized setup in which the participants collaboratively train and serve a large neural network, and where each participant only processes a subset of the model.…