2 papers
cs.LG2026
EFQ-Softmax: Exp-Free Quantization for Softmax
Haohui Han, Yuming Wan, Hongni Wang +4
Low-bit attention accelerates Transformer inference by moving the and matrix multiplications to FP8 or FP4 matrix engines. However, the softmax path often evaluates…
cs.LG2026
Compile to Compress: Boosting Formal Theorem Provers by Compiler Outputs
Guchan Li, Rui Tian, Hongning Wang
Large language models (LLMs) have demonstrated significant potential in formal theorem proving, yet state-of-the-art performance often necessitates prohibitive test-time compute vi…