Showing cs.ARShow all
2 papers · 1 filter
cs.AR2026
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference
Junyi Luo, Xinting Jiang, Tai-Hao Wen +9
Microscaling (MX) is now the standard for low-bit large language model (LLM) inference. Its 4-bit form MXFP4 still loses substantial accuracy, because existing MX formats fix eithe…
cs.AR2024
ConSmax: Hardware-Friendly Alternative Softmax with Learnable Parameters
Shiwei Liu, Guanchen Tao, Yifei Zou +7
The self-attention mechanism distinguishes transformer-based large language models (LLMs) apart from convolutional and recurrent neural networks. Despite the performance improvemen…