2 papers
cs.CL2026
Hidden Decoding at Scale: Latent Computation Scaling for Large Language Models
Aiwei Liu, Cheng Shi, Chuhan Wu +44
Scaling Large Language Models (LLMs) has been driven mainly by enlarging the Transformer backbone, but for an already-strong model this requires another round of costly pretraining…
cs.PF2024
DSO: A GPU Energy Efficiency Optimizer by Fusing Dynamic and Static Information
Qiang Wang, Laiyi Li, Weile Luo +2
Increased reliance on graphics processing units (GPUs) for high-intensity computing tasks raises challenges regarding energy consumption. To address this issue, dynamic voltage and…