3 papers
cs.LG2026
eOptShrinkQ: Near-Lossless KV Cache Compression Through Optimal Spectral Denoising and Quantization
Pei-Chun Su
We show that the key-value (KV) cache in transformer attention heads admits a natural decomposition into a low-rank \emph{shared context} component and a full-rank \emph{per-token}…
math.NA2026
Extracting Dual Analytic Geometries of Linear Transformations to Achieve Efficient Computation
Pei-Chun Su, Ronald R. Coifman
We propose a novel framework for fast integral operations by uncovering hidden geometries in the row and column structures of the underlying operators. This is accomplished through…
math.NA2025
Intrinsic and Extrinsic Organized Attention: Softmax Invariance and Network Sparsity
Oluwadamilola Fasina, Ruben V. C. Pohle, Pei-Chun Su +1
We examine the intrinsic (within the attention head) and extrinsic (amongst the attention heads) structure of the self-attention mechanism in transformers. Theoretical evidence for…