3 papers
cs.DC2026
RATrain: A Resource-Aware Training Runtime for Large Language Models on Bandwidth-Constrained Heterogeneous Supercomputing Platforms
Yao Lu, Shiqing Ma, Zhongzhi Luan +5
Production heterogeneous supercomputing platforms are increasingly used to host large language model (LLM) training workloads. However, existing GPU-oriented training runtimes typi…
cs.DC2026
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
Yao Lu, Zhongzhi Luan, Gen Li +6
Large language model (LLM) inference is limited by high computational cost and memory bandwidth demands, making deployment on heterogeneous many-core processors challenging. Taking…
cs.LG2024
Quantum Machine Learning in Log-based Anomaly Detection: Challenges and Opportunities
Jiaxing Qi, Chang Zeng, Zhongzhi Luan +6
Log-based anomaly detection (LogAD) is the main component of Artificial Intelligence for IT Operations (AIOps), which can detect anomalous that occur during the system on-the-fly.…