2 papers
cs.NI2026
Rethinking Network Topologies for Cost-Effective Mixture-of-Experts LLM Serving
Junsun Choi, Sam Son, Sunjin Choi +5
Mixture-of-experts (MoE) architectures have turned LLM serving into a cluster-scale workload in which communication consumes a considerable portion of LLM serving runtime. This has…
cs.AR2025
Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy Efficiency
Hansung Kim, Ruohan Richard Yan, Joshua You +2
Modern GPUs incorporate specialized matrix units such as Tensor Cores to accelerate GEMM operations, which are central to deep learning workloads. However, existing matrix unit des…