3 papers
cs.SE2026
From Noisy Telemetry to Actionable Warnings: GPU Failure Prediction in Industrial Clusters
Yongqian Sun, Run Zhu, Wenwei Gu +9
GPU clusters are critical infrastructure for AI services, but accurate and actionable GPU failure prediction remains a problem in production settings. We study ticket-linked teleme…
cs.CL2025
GTA: Grouped-head latenT Attention
Luoyang Sun, Cheng Deng, Jiwen Jiang +5
Attention mechanisms underpin the success of large language models (LLMs), yet their substantial computational and memory overhead poses challenges for optimizing efficiency and pe…
cs.CL2025
PLM: Efficient Peripheral Language Models Hardware-Co-Designed for Ubiquitous Computing
Cheng Deng, Luoyang Sun, Jiwen Jiang +10
While scaling laws have been continuously validated in large language models (LLMs) with increasing model parameters, the inherent tension between the inference demands of LLMs and…