4 papers
From Passive Response to Proactive Correction: Enhancing LLM Robustness Against Input Fact Perturbations
Ping Wang, Xiangguo Sun, Bingbing Xu +2
Large language models (LLMs) frequently produce confident yet factually incorrect responses when user inputs contain misleading premises, a phenomenon we attribute to fact perturba…
CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs
Qiuyi Qi, Jinjian Zhang, Mutian Bao +9
Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their r…
Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model
Pengxiang Cai, Tianchen Fang, Xiaohan Li +3
Reinforcement learning with verifiable rewards (RLVR) is widely viewed as a promising path toward continuously improving large language models. Recent works, however, suggest that…
From Misleading Queries to Accurate Answers: A Three-Stage Fine-Tuning Method for LLMs
Guocong Li, Weize Liu, Yihang Wu +4
Large language models (LLMs) exhibit excellent performance in natural language processing (NLP), but remain highly sensitive to the quality of input queries, especially when these…