2 papers
cs.AR2025
Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
Hannah Atmer, Yuan Yao, Thiemo Voigt +1
Energy consumption dictates the cost and environmental impact of deploying Large Language Models. This paper investigates the impact of on-chip SRAM size and operating frequency on…
eess.SY2025
Efficient CNN Inference on Ultra-Low-Power MCUs via Saturation-Aware Convolution
Shiming Li, Luca Mottola, Yuan Yao +1
Quantized CNN inference on ultra-low-power MCUs incurs unnecessary computations in neurons that produce saturated output values. These values are too extreme and are eventually cla…