3 papers
cs.DC2026
RATrain: A Resource-Aware Training Runtime for Large Language Models on Bandwidth-Constrained Heterogeneous Supercomputing Platforms
Yao Lu, Shiqing Ma, Zhongzhi Luan +5
Production heterogeneous supercomputing platforms are increasingly used to host large language model (LLM) training workloads. However, existing GPU-oriented training runtimes typi…
cs.DC2026
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
Yao Lu, Zhongzhi Luan, Gen Li +6
Large language model (LLM) inference is limited by high computational cost and memory bandwidth demands, making deployment on heterogeneous many-core processors challenging. Taking…
cs.SE2025
Beyond Window-Based Detection: A Graph-Centric Framework for Discrete Log Anomaly Detection
Jiaxing Qi, Chang Zeng, Zhongzhi Luan +5
Detecting anomalies in discrete event logs is critical for ensuring system reliability, security, and efficiency. Traditional window-based methods for log anomaly detection often s…