4 papers · 1 filter
ASTRA-sim 3.0: Next-Level Distributed Machine Learning Simulations via High-Fidelity GPU and Infrastructure Modeling
William Won, Jinsun Yoo, Tuan Ta +16
Distributed machine learning (ML) is a key paradigm for today's large-scale artificial intelligence applications. As model inference arises as an important use case, faithful model…
Scalable Synthesis of distributed LLM workloads through Symbolic Tensor Graphs
Changhai Man, Joongun Park, Hanjiang Wu +3
Optimizing the performance of large language models (LLMs) on large-scale AI training and inference systems requires a scalable and expressive mechanism to model distributed worklo…
Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
Seokjin Go, Joongun Park, Spandan More +5
The rapid scaling of Large Language Models (LLMs) has pushed training workloads far beyond the limits of single-node analysis, demanding a deeper understanding of how these models…
Enhancing Scalability and Performance in Influence Maximization with Optimized Parallel Processing
Hanjiang Wu, Huan Xu, Joongun Park +5
Influence Maximization (IM) is vital in viral marketing and biological network analysis for identifying key influencers. Given its NP-hard nature, approximate solutions are employe…