3 papers
cs.LG2026
How Modular Is a Frontier Mixture-of-Experts? A Pre-registered Causal Test in Which Apparent Expert Modularity Mostly Dissolves
Tony Salomone, Deep Gandhi, Ali Asaria
Sparse Mixture-of-Experts (MoE) models route each token to a few of many experts, inviting the hypothesis that experts form functional modules tied to capabilities or languages. We…
cs.AI2026
RobustDebias: Debiasing Language Models using Distributionally Robust Optimization
Deep Gandhi, Katyani Singh, Nidhi Hegde
Pretrained language models have been shown to exhibit biases and social stereotypes. Prior work on debiasing these models has largely focused on modifying embedding spaces during p…
cs.LG2022
A Federated Approach to Predicting Emojis in Hindi Tweets
Deep Gandhi, Jash Mehta, Nirali Parekh +3
The use of emojis affords a visual modality to, often private, textual communication. The task of predicting emojis however provides a challenge for machine learning as emoji use t…