2 papers
cs.DC2026
NestPipe: Large-Scale Recommendation Training on 1,500+ Accelerators via Nested Pipelining
Zhida Jiang, Zhaolong Xing, Huichao Chai +12
Modern recommendation models have increased to trillions of parameters. As cluster scales expand to O(1k), distributed training bottlenecks shift from computation and memory to dat…
cs.DC2024
LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
Yi Xiong, Hao Wu, Changxu Shao +6
The expanding context windows in large language models (LLMs) have greatly enhanced their capabilities in various applications, but they also introduce significant challenges in ma…