large language models 2adaptive pruning 1batch processing 1efficiency 1inference acceleration 1long-context inference 1proxy-kernel co-design 1sparse attention 1speculative decoding 1
From the 2 of 8 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
FOCUS: FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling
Xianglong Yan, Hong Liu, Chengzhu Bao +4
Large language models (LLMs) achieve remarkable performance but are expensive to deploy due to their enormous size. FP4 quantization, with formats such as MXFP4 and NVFP4, offers a…
cs.AI2024
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
Hanlin Tang, Yifu Sun, Decheng Wu +3
Large language models (LLMs) have proven to be very superior to conventional methods in various tasks. However, their expensive computations and high memory requirements are prohib…