8 citations · 8 across the 2 of their papers we have counts for
2 papers
cs.DC2026
A Scalable Recipe on SuperMUC-NG Phase 2: Efficient Large-Scale Training of Language Models
Ajay Navilarekal Rajgopal, Nikolai Solmsdorf
Large Language Models (LLMs) continue to demonstrate superior performance with increasing scale, yet training models with billions to trillions of parameters requires staggering co…
cs.LG2021★ 8 cited
Finding hidden-feature depending laws inside a data set and classifying it using Neural Network
Thilo Moshagen, Nihal Acharya Adde, Ajay Navilarekal Rajgopal
The logcosh loss function for neural networks has been developed to combine the advantage of the absolute error loss function of not overweighting outliers with the advantage of th…