papers

Publications (18)

cs.LG2026

Factored Gossip DiLoCo: Reducing Blocking Communication in DiLoCo

Chamin Hewa Koneputugodage, Thalaiyasingam Ajanthan, Sameera Ramasinghe +7

To make large-scale distributed training practical outside high-bandwidth datacenters, we must reduce blocking, high-volume synchronization. While DiLoCo communicates infrequently,…

cs.LG2026

Taming Curvature: Architecture Warm-Up for Stable Transformer Training

Sameera Ramasinghe, Ajanthan Thalaiyasingam, Hadi Mohaghegh Dolatabadi +6

Training billion-parameter Transformers is often brittle, with transient loss spikes and divergence that waste compute. Even though the recently developed Edge of Stability (EoS) t…

nucl-th2023

Effective field theory analysis of the Coulomb breakup of the one-neutron halo nucleus 19C

Pierre Capel, Daniel R. Phillips, Andrew Andis +29

We analyse the Coulomb breakup of 19C measured at 67A MeV at RIKEN. We use the Coulomb-Corrected Eikonal (CCE) approximation to model the reaction and describe the one-neutron halo…

physics.ins-det2015

Low energy neutron background in deep underground laboratories

Andreas Best, Joachim Gorres, Matthias Junker +6

The natural neutron background influences the maximum achievable sensitivity in most deep underground nuclear, astroparticle and double-beta decay physics experiments. Reliable neu…

cs.LG2026

NuMuon: Nuclear-Norm-Constrained Muon for Compressible LLM Training

Hadi Mohaghegh Dolatabadi, Thalaiyasingam Ajanthan, Sameera Ramasinghe +7

The rapid progress of large language models (LLMs) is increasingly constrained by memory and deployment costs, motivating compression methods for practical deployment. Many state-o…

cs.LG2026

Unextractable Protocol Models: Collaborative Training and Inference without Weight Materialization

Alexander Long, Chamin Hewa Koneputugodage, Thalaiyasingam Ajanthan +5

We consider a decentralized setup in which the participants collaboratively train and serve a large neural network, and where each participant only processes a subset of the model.…