2 papers
cs.LG2025
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
Yifang Chen, Xiaoyu Li, Yingyu Liang +3
The key-value (KV) cache in the tensor version of transformers presents a significant bottleneck during inference. While previous work analyzes the fundamental space complexity bar…
cs.LG2025
Theoretical Guarantees for High Order Trajectory Refinement in Generative Flows
Chengyue Gong, Xiaoyu Li, Yingyu Liang +4
Flow matching has emerged as a powerful framework for generative modeling, offering computational advantages over diffusion models by leveraging deterministic Ordinary Differential…