2 papers
cs.DC2025
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
Zizhao Mo, Jianxiong Liao, Huanle Xu +2
The significant resource demands in LLM serving prompts production clusters to fully utilize heterogeneous hardware by partitioning LLM models across a mix of high-end and low-end…
cs.AR2025
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
Huanyu Qu, Weihao Zhang, Junfeng Lin +4
To efficiently support large-scale NNs, multi-level hardware, leveraging advanced integration and interconnection technologies, has emerged as a promising solution to counter the s…