3 papers
cs.LG2025
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
Tianyu Fu, Haofeng Huang, Xuefei Ning +10
Sliding-window attention offers a hardware-efficient solution to the memory and throughput challenges of Large Language Models (LLMs) in long-context scenarios. Existing methods ty…
cs.LG2024
Canvas: End-to-End Kernel Architecture Search in Neural Networks
Chenggang Zhao, Genghan Zhang, Ao Shen +1
The demands for higher performance and accuracy in neural networks (NNs) never end. Existing tensor compilation and Neural Architecture Search (NAS) techniques orthogonally optimiz…
cs.LG2024
CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models
Donghyun Lee, Je-Yong Lee, Genghan Zhang +2
Large Language Models (LLMs) have dramatically advanced AI applications, yet their deployment remains challenging due to their immense inference costs. Recent studies ameliorate th…