6 papers
A Novel Block-Alternating Iterative Algorithm for Retrieving Top- Elements from Factorized Tensors
Chuanfu Xiao, Jiaxin Zeng
Tensors, especially higher-order tensors, are typically represented in low-rank formats to preserve the main information of the high-dimensional data while saving memory space. In…
HaTT: Hadamard avoiding TT recompression
Zhonghao Sun, Jizu Huang, Chuanfu Xiao +1
The Hadamard product of tensor train (TT) tensors is a fundamental nonlinear operation in scientific computing and data analysis. However, due to its tendency to significantly incr…
AnTKV: Anchor Token-Aware Sub-Bit Vector Quantization for KV Cache in Large Language Models
Zeyu Li, Chuanfu Xiao, Yang Wang +6
Quantization has emerged as an effective and lightweight solution to reduce the memory footprint of the KV cache in Large Language Models. Nevertheless, minimizing the accuracy deg…
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
Qianchao Zhu, Jiangfei Duan, Chang Chen +6
Large language models (LLMs) now support extremely long context windows, but the quadratic complexity of vanilla attention results in significantly long Time-to-First-Token (TTFT)…
Provable Low-Rank Tensor-Train Approximations in the Inverse of Large-Scale Structured Matrices
Chuanfu Xiao, Kejun Tang, Zhitao Zhu
This paper studies the low-rank property of the inverse of a class of large-scale structured matrices in the tensor-train (TT) format, which is typically discretized from different…
APTT: An accuracy-preserved tensor-train method for the Boltzmann-BGK equation
Zhitao Zhu, Chuanfu Xiao, Kejun Tang +2
Solving the Boltzmann-BGK equation with traditional numerical methods suffers from high computational and memory costs due to the curse of dimensionality. In this paper, we propose…