Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
DCPO: Dynamic Clipping Policy Optimization
Shihui Yang, Chengfeng Dou, Peidong Guo +4
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising framework for enhancing the reasoning capabilities of large language models. However, existing appr…
cs.CL2025
Baichuan 2: Open Large-scale Language Models
Aiyuan Yang, Bin Xiao, Bingning Wang +52
Large language models (LLMs) have demonstrated remarkable performance on a variety of natural language tasks based on just a few examples of natural language instructions, reducing…