10 citations · 10 across the 2 of their papers we have counts for
2 papers
cs.CE2024
Task Scheduling for Efficient Inference of Large Language Models on Single Moderate GPU Systems
Wenxiang Lin, Xinglin Pan, Shaohuai Shi +2
Large language models~(LLMs) are known for their high demand on computing resources and memory due to their substantial model size, which leads to inefficient inference on moderate…
cs.DC2023★ 10 cited
FusionAI: Decentralized Training and Deploying LLMs with Massive Consumer-Level GPUs
Zhenheng Tang, Yuxin Wang, Xin He +8
The rapid growth of memory and computation requirements of large language models (LLMs) has outpaced the development of hardware, hindering people who lack large-scale high-end GPU…