12 citations · 28 across the 15 of their papers we have counts for
Showing cs.PFShow all
2 papers · 1 filter
cs.PF2026
Compass: Dissecting Communication and Computation Operators for Efficient LLM Training
Guangyu Xiang, Lin Zhang, Haoxuan Yu +3
Overlapping communication and computation operators is a common practice to hide communication overheads, accelerating large language models (LLMs) training on GPU clusters. Existi…
cs.PF2023
Dissecting the Runtime Performance of the Training, Fine-tuning, and Inference of Large Language Models
Longteng Zhang, Xiang Liu, Zeyu Li +8
Large Language Models (LLMs) have seen great advance in both academia and industry, and their popularity results in numerous open-source frameworks and techniques in accelerating L…