2 papers
cs.LG2026
QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization
Yipin Guo, Arun M George, Jie Fu +3
Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. However, existing PTQ methods often fail to ge…
cs.LG2025
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
Patrick Yubeaton, Tareq Mahmoud, Shehab Naga +8
As they become more capable, large language models (LLMs) have continued to rapidly increase in size. This has exacerbated the difficulty in running state of the art LLMs on small,…