Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Decoupling Understanding from Reasoning via Problem Space Mapping for Small-Scale Model Reasoning
Li Wang, Changhao Zhang, Zengqi Xiu +4
Despite recent advances in the reasoning capabilities of Large Language Models (LLMs), improving the reasoning ability of Small Language Models (SLMs, e.g., up to 1.5B parameters)…
cs.CL2025
DCPO: Dynamic Clipping Policy Optimization
Shihui Yang, Chengfeng Dou, Peidong Guo +4
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising framework for enhancing the reasoning capabilities of large language models. However, existing appr…