4 papers
A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
William Nixon, Jon Durbin, Florian Standhartinger +2
Large Language Model (LLM) serving has become a critical cloud workload, and realistic traces are essential for motivating and benchmarking serving systems. However, existing LLM s…
Understanding and Detecting Scalability Faults in Large-Scale Distributed Systems
Hao-Nan Zhu, Goodness Ayinmode, Cesar A. Stuardo +2
Scalable distributed systems form the backbone of modern computing infrastructure. However, as scale grows, system complexity may lead to scalability faults. Scalability faults are…
StorRep: Storage Research Experiment Patterns on Chameleon Cloud and Trovi
Ray A. O. Sinurat, Yuyang Huang, Nanqinqin Li +4
Storage experiments are vital to advancing storage research, but creating extensible and reproducible storage artifacts can be a challenging task. Our research has shown that only…
Alchemist: Towards the Design of Efficient Online Continual Learning System
Yuyang Huang, Yuhan Liu, Haryadi S. Gunawi +2
Continual learning has become a promising solution to refine large language models incrementally by leveraging user feedback. In particular, online continual learning - iteratively…