1 paper
Bing Tian, Haikun Liu, Xiaocheng Zhong +5
Block-wise diffusion large language models (dLLMs) decode sequentially at the block level, enabling effective KV-cache reuse across blocks but making inter-block decoding strictly…