4 papers
Distributed Hybrid Parallelism for Large Language Models: Comparative Study and System Design Guide
Hossam Amer, Rezaul Karim, Ali Pourranjbar +3
With the rapid growth of large language models (LLMs), a wide range of methods have been developed to distribute computation and memory across hardware devices for efficient traini…
A Distributed Generative AI Approach for Heterogeneous Multi-Domain Environments under Data Sharing constraints
Youssef Tawfilis, Hossam Amer, Minar El-Aasser +1
Federated Learning has gained attention for its ability to enable multiple nodes to collaboratively train machine learning models without sharing raw data. At the same time, Genera…
FLOP-Efficient Training: Early Stopping Based on Test-Time Compute Awareness
Hossam Amer, Maryam Dialameh, Hossein Rajabzadeh +3
Scaling training compute, measured in FLOPs, has long been shown to improve the accuracy of large language models, yet training remains resource-intensive. Prior work shows that in…
Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection
Mohammad Mahdi Moradi, Hossam Amer, Sudhir Mudur +3
Learning to adapt pretrained language models to unlabeled, out-of-distribution data is a critical challenge, as models often falter on structurally novel reasoning tasks even while…