4 papers · 1 filter
DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding
Yanhua Jiao, Tianyi Wu, Xiaoxi Sun +6
While parallel decoding is central to the efficiency of Diffusion Large Language Models (dLLMs), current strategies are often hindered by overly conservative confidence thresholds.…
: Unlocking LLM Reasoning via Reinforcement Learning with Re-solving
Pinzheng Wang, Shuli Xu, Juntao Li +4
Reinforcement learning with verifiable rewards (RLVR) has shown promise in enhancing the reasoning performance of large language models (LLMs) by increasing test-time compute. Howe…
Improving Value-based Process Verifier via Low-Cost Variance Reduction
Zetian Sun, Dongfang Li, Baotian Hu +1
Large language models (LLMs) have achieved remarkable success in a wide range of tasks. However, their reasoning capabilities, particularly in complex domains like mathematics, rem…
Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?
Zetian Sun, Dongfang Li, Xuhui Chen +2
The alignment of language models~(LMs) with human preferences is critical for building reliable AI systems. The problem is typically framed as optimizing an LM policy to maximize t…