125 citations · 154 across the 10 of their papers we have counts for
Showing cs.ARShow all
2 papers · 1 filter
cs.AR2025
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
Coleman Hooper, Charbel Sakr, Ben Keller +4
Quantization is a powerful tool to improve large language model (LLM) inference efficiency by utilizing more energy-efficient low-precision datapaths and reducing memory footprint.…
cs.AR2021
Softermax: Hardware/Software Co-Design of an Efficient Softmax for Transformers
Jacob R. Stevens, Rangharajan Venkatesan, Steve Dai +2
Transformers have transformed the field of natural language processing. This performance is largely attributed to the use of stacked self-attention layers, each of which consists o…