Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
FastCache: Fast Caching for Diffusion Transformer Through Learnable Linear Approximation
Dong Liu, Yanxuan Yu, Jiayi Zhang +3
Diffusion Transformers (DiT) are powerful generative models but remain computationally intensive due to their iterative structure and deep transformer stacks. To alleviate this ine…
cs.LG2026
MKA: Memory-Keyed Attention for Efficient Long-Context Reasoning
Dong Liu, Yanxuan Yu, Ben Lengerich +1
As long-context language modeling becomes increasingly important, the cost of maintaining and attending to large Key/Value (KV) caches grows rapidly, becoming a major bottleneck in…