Showing cs.ARShow all
2 papers · 1 filter
cs.AR2026
SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference
Aradhana Mohan Parvathy, Soumendu Kumar Ghosh, Shamik Kundu +4
The rapid growth in sizes of Large language models (LLMs) results in high compute and memory costs during inference. Quantization has been a significant pathway to addressing this…
cs.AR2025
StruM: Structured Mixed Precision for Efficient Deep Learning Hardware Codesign
Michael Wu, Arnab Raha, Deepak A. Mathaikutty +3
In this paper, we propose StruM, a novel structured mixed-precision-based deep learning inference method, co-designed with its associated hardware accelerator (DPU), to address the…