1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.LG2025★ 1 cited
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
Yifang Chen, Xiaoyu Li, Yingyu Liang +3
The key-value (KV) cache in the tensor version of transformers presents a significant bottleneck during inference. While previous work analyzes the fundamental space complexity bar…
cs.LG2024
Theoretical Constraints on the Expressive Power of -based Tensor Attention Transformers
Xiaoyu Li, Yingyu Liang, Zhenmei Shi +2
Tensor Attention extends traditional attention mechanisms by capturing high-order correlations across multiple modalities, addressing the limitations of classical matrix-based atte…
cs.LG2024
Circuit Complexity Bounds for RoPE-based Transformer Architecture
Bo Chen, Xiaoyu Li, Yingyu Liang +3
Characterizing the express power of the Transformer architecture is critical to understanding its capacity limits and scaling law. Recent works provide the circuit complexity bound…