2 papers
cs.DC2026
W4A16 Mixed-Precision Matrix Multiplication on Decoupled Architecture: Kernel Design and Memory Bottleneck Analysis for Ascend NPUs
Yuanhong He, Peiyu Niu, Jun Chen +2
As Large Language Models (LLMs) scale, weight-only quantization (W4A16: 4-bit weights, 16-bit activations) becomes critical for reducing memory footprint with minimal accuracy loss…
cs.LG2025
LeForecast: Enterprise Hybrid Forecast by Time Series Intelligence
Zheng Tan, Yiwen Nie, Wenfa Wu +22
Demand is spiking in industrial fields for multidisciplinary forecasting, where a broad spectrum of sectors needs planning and forecasts to streamline intelligent business manageme…