2 papers
cs.CL2025
MoxE: Mixture of xLSTM Experts with Entropy-Aware Routing for Efficient Language Modeling
Abdoul Majid O. Thiombiano, Brahim Hnich, Ali Ben Mrad +1
This paper introduces MoxE, a novel architecture that synergistically combines the Extended Long Short-Term Memory (xLSTM) with the Mixture of Experts (MoE) framework to address cr…
cs.LG2025
Distil-xLSTM: Learning Attention Mechanisms through Recurrent Structures
Abdoul Majid O. Thiombiano, Brahim Hnich, Ali Ben Mrad +1
The current era of Natural Language Processing (NLP) is dominated by Transformer models. However, novel architectures relying on recurrent mechanisms, such as xLSTM and Mamba, have…