1 citations · 1 across the 3 of their papers we have counts for
12 papers
Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning
Jiaxing Qi, Zhongzhi Luan, Hongyu Zhang +5
Large language models (LLMs) are increasingly used to interpret operational evidence and assist incident response in cloud-native microservice systems. However, recovery-oriented u…
FDLoRA: Personalized Federated Learning of Large Language Model via Dual LoRA Tuning
Yao Lu, Jiaxing QI, Zhongzhi Luan +4
Large language models (LLMs) have emerged as important components across various fields, yet their training requires substantial computation resources and abundant labeled data. It…
xGR: Efficient Generative Recommendation Serving at Scale
Qingxiao Sun, Tongxuan Liu, Shen Zhang +13
Recommendation system delivers substantial economic benefits by providing personalized predictions. Generative recommendation (GR) integrates LLMs to enhance the understanding of l…
RATrain: A Resource-Aware Training Runtime for Large Language Models on Bandwidth-Constrained Heterogeneous Supercomputing Platforms
Yao Lu, Shiqing Ma, Zhongzhi Luan +5
Production heterogeneous supercomputing platforms are increasingly used to host large language model (LLM) training workloads. However, existing GPU-oriented training runtimes typi…
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
Yao Lu, Zhongzhi Luan, Gen Li +6
Large language model (LLM) inference is limited by high computational cost and memory bandwidth demands, making deployment on heterogeneous many-core processors challenging. Taking…
CB-SpMV:A Data Aggregating and Balance Algorithm for Cache-Friendly Block-Based SpMV on GPUs
Xing Cong, Fukai Sun, Yifan Chen +3
Sparse matrix-vector multiplication (SpMV) is crucial in computational science, engineering, and machine learning. Despite substantial efforts to improve SpMV performance on GPUs t…