Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
Liang Chen, Xueting Han, Li Shen +2
Supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR) are two widely used post-training paradigms for improving the reasoning ability of large lang…
cs.CL2026
EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget
Liang Chen, Xueting Han, Qizhou Wang +4
Balancing exploration and exploitation remains a central challenge in reinforcement learning with verifiable rewards (RLVR) for large language models (LLMs). Current RLVR methods o…
cs.CL2025
A Survey of the Evolution of Language Model-Based Dialogue Systems: Data, Task and Models
Hongru Wang, Lingzhi Wang, Yiming Du +4
Dialogue systems (DS), including the task-oriented dialogue system (TOD) and the open-domain dialogue system (ODD), have always been a fundamental task in natural language processi…