5 citations · 7 across the 3 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2024★ 1 cited
PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
Dongjie Yang, XiaoDong Han, Yan Gao +3
Large Language Models (LLMs) have shown remarkable comprehension abilities but face challenges in GPU memory usage during inference, hindering their scalability for real-time appli…
cs.CL2023★ 2 cited
Linearized Relative Positional Encoding
Zhen Qin, Weixuan Sun, Kaiyue Lu +6
Relative positional encoding is widely used in vanilla and linear transformers to represent positional information. However, existing encoding methods of a vanilla transformer are…
cs.CL2023★ 5 cited
Toeplitz Neural Network for Sequence Modeling
Zhen Qin, Xiaodong Han, Weixuan Sun +6
Sequence modeling has important applications in natural language processing and computer vision. Recently, the transformer-based models have shown strong performance on various seq…