Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Reversible Diffusion Decoding for Diffusion Language Models
Xinyun Wang, Min Zhang, Sen Cui +4
Diffusion language models enable parallel token generation through block-wise decoding, but their irreversible commitments can lead to stagnation, where the reverse diffusion proce…
cs.CL2025
CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences
Ziran Qin, Yuchen Cao, Mingbao Lin +5
Large language models (LLMs) excel at processing long sequences, boosting demand for key-value (KV) caching. While recent efforts to evict KV cache have alleviated the inference bu…