activity
20202025
most citedScalable and Efficient MoE Training for Multitask Multilingual Models

33 citations · 33 across the 4 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2025

Privacy-Preserving Transformers: SwiftKey's Differential Privacy Implementation

Abdelrahman Abouelenin, Mohamed Abdelrehim, Raffy Fahim +2

In this paper we train a transformer using differential privacy (DP) for language modeling in SwiftKey. We run multiple experiments to balance the trade-off between the model size,…

cs.CL20257 cited

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Microsoft, :, Abdelrahman Abouelenin +73

We introduce Phi-4-Mini and Phi-4-Multimodal, compact yet highly capable language and multimodal models. Phi-4-Mini is a 3.8-billion-parameter language model trained on high-qualit…

cs.CL2022

Domain Specific Sub-network for Multi-Domain Neural Machine Translation

Amr Hendy, Mohamed Abdelghaffar, Mohamed Afify +1

This paper presents Domain-Specific Sub-network (DoSS). It uses a set of masks obtained through pruning to define a sub-network for each domain and finetunes the sub-network parame…

cs.CL202133 cited

Scalable and Efficient MoE Training for Multitask Multilingual Models

Young Jin Kim, Ammar Ahmad Awan, Alexandre Muzio +6

The Mixture of Experts (MoE) models are an emerging class of sparsely activated deep learning models that have sublinear compute costs with respect to their parameters. In contrast…

cs.CL2020

Score Combination for Improved Parallel Corpus Filtering for Low Resource Conditions

Muhammad N. ElNokrashy, Amr Hendy, Mohamed Abdelghaffar +3

This paper describes our submission to the WMT20 sentence filtering task. We combine scores from (1) a custom LASER built for each source language, (2) a classifier built to distin…