Showing cs.DCShow all
3 papers · 1 filter
cs.DC2026
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
Hongyu Chen, Letian Ruan, Zilin Xu +6
LoRA enables efficient customization of LLMs and is widely used in multi-tenant and multi-task serving. However, emerging model architectures such as MoE significantly increase LoR…
cs.DC2026
CCCL: Node-Spanning GPU Collectives with CXL Memory Pooling
Dong Xu, Han Meng, Xinyu Chen +11
Large language models (LLMs) training or inference across multiple nodes introduces significant pressure on GPU memory and interconnect bandwidth. The Compute Express Link (CXL) sh…
cs.DC2025
Rethinking Dynamic Networks and Heterogeneous Computing with Automatic Parallelization
Ruilong Wu, Xinjiao Li, Yisu Wang +2
Hybrid parallelism techniques are essential for efficiently training large language models (LLMs). Nevertheless, current automatic parallel planning frameworks often overlook the s…