3 papers
cs.DC2026
ARGUS: Production-Scale Tracing and Performance Diagnosis for over 10,000-GPU Clusters
Jiasheng Zhou, Longbin Zeng, Clavis Chen +5
Large-scale LLM training requires always-on, fine-grained observability for effective performance diagnosis at scale. Coarse resource monitors alone cannot localize root causes, an…
cs.PF2025
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
Zizhuo Fu, Xiaotian Guo, Wenxuan Zeng +6
Large language models (LLMs) have demonstrated remarkable proficiency in a wide range of natural language processing applications. However, the high energy and latency overhead ind…
cs.CL2025
Orchestrating Dual-Boundaries: An Arithmetic Intensity Inspired Acceleration Framework for Diffusion Language Models
Linye Wei, Wenjue Chen, Pingzhi Tang +4
Diffusion-based large language models (dLLMs) have recently gained significant attention for their exceptional performance and inherent potential for parallel decoding. Existing fr…