3 papers
cs.LG2026
Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models
Ali Janati, Kaoutar El Maghraoui, Xinyi Luo +3
Mixture-of-Experts (MoE) models decouple total parameters from per-token compute, but deployment still requires storing every expert. Recent theory shows that pruning experts with…
cs.AI2026
Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers
Ali Janati, Kaoutar El Maghraoui, Andrei Kanavalau +1
Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition faster, but its solutions do not hold. All nine configurations on…
cs.AI2026
Uncertainty-Aware Multimodal Emotion Recognition through Dirichlet Parameterization
Rémi Grzeczkowicz, Eric Soriano, Ali Janati +4
In this work, we present a lightweight and privacy-preserving Multimodal Emotion Recognition (MER) framework designed for deployment on edge devices. To demonstrate framework's ver…