14 citations · 41 across the 14 of their papers we have counts for
4 papers · 1 filter
FusionLLM: A Decentralized LLM Training System on Geo-distributed GPUs with Adaptive Compression
Zhenheng Tang, Xueze Kang, Yiming Yin +11
To alleviate hardware scarcity in training large deep neural networks (DNNs), particularly large language models (LLMs), we present FusionLLM, a decentralized training system desig…
FusionAI: Decentralized Training and Deploying LLMs with Massive Consumer-Level GPUs
Zhenheng Tang, Yuxin Wang, Xin He +8
The rapid growth of memory and computation requirements of large language models (LLMs) has outpaced the development of hardware, hindering people who lack large-scale high-end GPU…
Benchmarking the Performance and Energy Efficiency of AI Accelerators for AI Training
Yuxin Wang, Qiang Wang, Shaohuai Shi +4
Deep learning has become widely used in complex AI applications. Yet, training a deep neural network (DNNs) model requires a considerable amount of calculations, long running time,…
A Distributed Synchronous SGD Algorithm with Global Top- Sparsification for Low Bandwidth Networks
Shaohuai Shi, Qiang Wang, Kaiyong Zhao +4
Distributed synchronous stochastic gradient descent (S-SGD) has been widely used in training large-scale deep neural networks (DNNs), but it typically requires very high communicat…