2 papers
cs.LG2026
Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking
Ali Saheb Pasand, Elvis Dohmatob
Grokking is the phenomenon whereby, unlike the training performance, which peaks early in the training process, the test/generalization performance of a model stagnates over arbitr…
cs.AI2026
REAM: Merging Improves Pruning of Experts in LLMs
Saurav Jha, Maryam Hashemzadeh, Ali Saheb Pasand +3
Mixture-of-Experts (MoE) large language models (LLMs) are among the top-performing architectures. The largest models, often with hundreds of billions of parameters, pose significan…