collaborators

6 papers

math.NA2025

A Novel Block-Alternating Iterative Algorithm for Retrieving Top- Elements from Factorized Tensors

Chuanfu Xiao, Jiaxin Zeng

Tensors, especially higher-order tensors, are typically represented in low-rank formats to preserve the main information of the high-dimensional data while saving memory space. In…

math.NA2025

HaTT: Hadamard avoiding TT recompression

Zhonghao Sun, Jizu Huang, Chuanfu Xiao +1

The Hadamard product of tensor train (TT) tensors is a fundamental nonlinear operation in scientific computing and data analysis. However, due to its tendency to significantly incr…

cs.CL2025

AnTKV: Anchor Token-Aware Sub-Bit Vector Quantization for KV Cache in Large Language Models

Zeyu Li, Chuanfu Xiao, Yang Wang +6

Quantization has emerged as an effective and lightweight solution to reduce the memory footprint of the KV cache in Large Language Models. Nevertheless, minimizing the accuracy deg…

cs.CL2025

SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention

Qianchao Zhu, Jiangfei Duan, Chang Chen +6

Large language models (LLMs) now support extremely long context windows, but the quadratic complexity of vanilla attention results in significantly long Time-to-First-Token (TTFT)…

math.NA2025

Provable Low-Rank Tensor-Train Approximations in the Inverse of Large-Scale Structured Matrices

Chuanfu Xiao, Kejun Tang, Zhitao Zhu

This paper studies the low-rank property of the inverse of a class of large-scale structured matrices in the tensor-train (TT) format, which is typically discretized from different…

math.NA2024

APTT: An accuracy-preserved tensor-train method for the Boltzmann-BGK equation

Zhitao Zhu, Chuanfu Xiao, Kejun Tang +2

Solving the Boltzmann-BGK equation with traditional numerical methods suffers from high computational and memory costs due to the curse of dimensionality. In this paper, we propose…