4 papers
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
Jiu Chen, Shuangyan Yang, Xu Xiong +4
Decentralized LLM inference distributes computation among heterogeneous nodes across the internet, offering a performant and cost-efficient solution, alternative to traditional cen…
Hybrid Adaptive Tuning for Tiered Memory Systems
Xi Wang, Jie Liu, Shuangyan Yang +3
Memory tiering provides a cost-effective solution to increase memory capacity, utilization, and even bandwidth. Memory tiering relies on system software for memory profiling, detec…
Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered Memory
Jie Ren, Bin Ma, Shuangyan Yang +4
Deep learning recommendation models (DLRMs) are widely used in industry, and their memory capacity requirements reach the terabyte scale. Tiered memory architectures provide a cost…
Tuning Fast Memory Size based on Modeling of Page Migration for Tiered Memory
Shangye Chen, Jin Huang, Shuangyan Yang +7
Tiered memory, built upon a combination of fast memory and slow memory, provides a cost-effective solution to meet ever-increasing requirements from emerging applications for large…