3 papers
cs.LG2026
MKA: Memory-Keyed Attention for Efficient Long-Context Reasoning
Dong Liu, Yanxuan Yu, Ben Lengerich +1
As long-context language modeling becomes increasingly important, the cost of maintaining and attending to large Key/Value (KV) caches grows rapidly, becoming a major bottleneck in…
cs.LG2025
FastCache: Fast Caching for Diffusion Transformer Through Learnable Linear Approximation
Dong Liu, Yanxuan Yu, Jiayi Zhang +3
Diffusion Transformers (DiT) are powerful generative models but remain computationally intensive due to their iterative structure and deep transformer stacks. To alleviate this ine…
cs.DC2024
Designing Large Foundation Models for Efficient Training and Inference: A Survey
Dong Liu, Yanxuan Yu, Yite Wang +5
This paper focuses on modern efficient training and inference technologies on foundation models and illustrates them from two perspectives: model and system design. Model and Syste…