large language models 2adaptive pruning 1batch processing 1efficiency 1inference acceleration 1long-context inference 1proxy-kernel co-design 1sparse attention 1speculative decoding 1
From the 2 of 9 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
FOCUS: FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling
Xianglong Yan, Hong Liu, Chengzhu Bao +4
Large language models (LLMs) achieve remarkable performance but are expensive to deploy due to their enormous size. FP4 quantization, with formats such as MXFP4 and NVFP4, offers a…
cs.AI2026
LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing
Wen Zan, Jiaqi Zhang, Jianchao Tan +11
DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remains constrained by the indexer's expensive…