collaborators

5 papers

cs.OS2026

CloakLM: Obfuscating GPU Memory Layout to Mitigate Model Ex-filtration for Serving

Kunal Jain, Seokjin Go, Divya Mahajan

Large foundation models deployed on third-party and shared accelerator infrastructure face a practical risk of model exfiltration that existing defenses do not fully address. In co…

cs.DC2026

ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving

Seokjin Go, Marko Scrbak, Ephrem Wu +2

In distributed Mixture-of-Experts (MoE) inference, input-dependent token routing interacts with GPU performance variability to create persistent stragglers under synchronized execu…

cs.DC2025

Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective

Seokjin Go, Joongun Park, Spandan More +5

The rapid scaling of Large Language Models (LLMs) has pushed training workloads far beyond the limits of single-node analysis, demanding a deeper understanding of how these models…

cs.DC2025

Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications

Seonho Lee, Jihwan Oh, Junkyum Kim +3

This paper provides an in-depth characterization of GPU-accelerated systems, to understand the interplay between overlapping computation and communication which is commonly employe…

cs.LG2025

MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing

Seokjin Go, Divya Mahajan

Mixture-of-Experts (MoE) model architecture has emerged as a promising solution for scaling transformer models efficiently, offering sparse activation that reduces computational co…