activity
20242026
most citedLearning Fine-grained Parameter Sharing via Sparse Tensor Decomposition

1 citations · 2 across the 4 of their papers we have counts for

collaborators

9 papers

cs.LG20261 cited

Learning Fine-grained Parameter Sharing via Sparse Tensor Decomposition

Cem Üyük, Mike Lasby, Mohamed Yassin +2

Large neural networks achieve state-of-the-art performance on many tasks, yet their sheer size hinders deployment on resource-constrained devices. Among existing compression approa…

cs.LG2026

SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse Training

Mohammed Adnan, Rohan Jain, Tom Jacobs +4

Dynamic Sparse Training (DST) methods train neural networks by maintaining sparsity while dynamically adapting the network topology. Despite the promise of reduced computation, DST…

cs.LG20261 cited

REAP the Experts: Why Pruning Prevails for One-Shot MoE compression

Mike Lasby, Ivan Lazarevich, Nish Sinnadurai +3

Sparsely-activated Mixture-of-Experts (SMoE) models offer efficient pre-training and low latency but their large parameter counts create significant memory overhead, motivating res…

cs.CL2026

Reducing Hallucinations in LLMs via Factuality-Aware Preference Learning

Sindhuja Chaduvula, Ahmed Y. Radwan, Azib Farooq +2

Preference alignment methods such as RLHF and Direct Preference Optimization (DPO) improve instruction following, but they can also reinforce hallucinations when preference judgmen…

cs.LG2025

Sparse Training from Random Initialization: Aligning Lottery Ticket Masks using Weight Symmetry

Mohammed Adnan, Rohan Jain, Ekansh Sharma +2

The Lottery Ticket Hypothesis (LTH) suggests there exists a sparse LTH mask and weights that achieve the same generalization performance as the dense model while using significantl…

cs.LG2025

What is Left After Distillation? How Knowledge Transfer Impacts Fairness and Bias

Aida Mohammadshahi, Yani Ioannou

Knowledge Distillation is a commonly used Deep Neural Network (DNN) compression method, which often maintains overall generalization performance. However, we show that even for bal…