4 papers
Self-Routing: Parameter-Free Expert Routing from Hidden States
Jama Hussein Mohamud, Drew Wagner, Mirco Ravanelli
Mixture-of-Experts (MoE) layers increase model capacity by activating only a small subset of experts per token, and typically rely on a learned router to map hidden states to exper…
Adaptive Order Policies for Masked Diffusion
Jama Hussein Mohamud, Mohsin Hasan, Mirco Ravanelli +1
Masked diffusion models have seen great success in capturing data distributions over discrete sequences in domains such as text and proteins. These models generate data by iterativ…
An empirical study of task and feature correlations in the reuse of pre-trained models
Jama Hussein Mohamud, Willie Brink
Pre-trained neural networks are commonly used and reused in the machine learning community. Alice trains a model for a particular task, and a part of her neural network is reused b…
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments
Yuran Li, Jama Hussein Mohamud, Chongren Sun +2
Large language models (LLMs) are being widely applied across various fields, but as tasks become more complex, evaluating their responses is increasingly challenging. Compared to h…