1 paper
Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen
This work studies how to adaptively recompute key-value (KV) caches for diffusion large language models (DLMs) to maximize prediction accuracy while minimizing decoding latency. Pr…