3 papers
cs.CL2026
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning
Zicheng Xu, Ruixuan Zhang, Yu-Neng Chuang +7
Large Language Models (LLMs) achieve remarkable reasoning capabilities through reinforcement learning (RL) post-training. However, existing RL post-training commonly relies on unif…
cs.AI2026
DTS: Enhancing Large Reasoning Models via Decoding Tree Sketching
Zicheng Xu, Xiuyi Lou, Guanchu Wang +6
Large Reasoning Models (LRMs) achieve remarkable inference-time improvements through parallel thinking. However, existing methods rely on redundant sampling of reasoning trajectori…
cs.CL2025
Self-ensemble: Mitigating Confidence Mis-calibration for Large Language Models
Zicheng Xu, Guanchu Wang, Guangyao Zheng +4
Although Large Language Models (LLMs) perform well in general fields, they exhibit a confidence distortion problem on multi-choice question-answering (MCQA), particularly as the nu…