2 papers
cs.AI2026
LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models
Xingru Chen, Zelang Liang, Yongjia Ma +4
Diffusion-based Large Language Models(DLLMs) enable parallel generation via Semi-Autoregressive (SAR) decoding in text generation. However, current methods suffer from severe opera…
cs.LG2026
HMPO: Hybrid Median-length Policy Optimization for Chain-of-Thought Compression
Minghui Zheng, Hongxu Chen, Huimin Ren +8
Large language models achieve remarkable performance via extended chain-of-thought (CoT) reasoning, yet this lengthy process incurs substantial inference overhead. Existing CoT com…