Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Improving Value-based Process Verifier via Low-Cost Variance Reduction
Zetian Sun, Dongfang Li, Baotian Hu +1
Large language models (LLMs) have achieved remarkable success in a wide range of tasks. However, their reasoning capabilities, particularly in complex domains like mathematics, rem…
cs.AI2026
Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?
Zetian Sun, Dongfang Li, Xuhui Chen +2
The alignment of language models~(LMs) with human preferences is critical for building reliable AI systems. The problem is typically framed as optimizing an LM policy to maximize t…