large language models 2adaptive pruning 1batch processing 1efficiency 1inference acceleration 1long-context inference 1proxy-kernel co-design 1sparse attention 1speculative decoding 1
From the 2 of 17 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
FOCUS: FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling
Xianglong Yan, Hong Liu, Chengzhu Bao +4
Large language models (LLMs) achieve remarkable performance but are expensive to deploy due to their enormous size. FP4 quantization, with formats such as MXFP4 and NVFP4, offers a…
cs.AI2025
SpecExit: Accelerating Large Reasoning Model via Speculative Exit
Rubing Yang, Huajun Bai, Song Liu +7
Despite their strong performance on reasoning tasks, large reasoning models (LRMs) often suffer from overthinking, producing unnecessarily long outputs and incurring high end-to-en…