2 papers
cs.CL2026
SSPO: Subsentence-level Policy Optimization
Kun Yang, Zikang chen, Yanmeng Wang +4
As a key component of large language model (LLM) post-training, Reinforcement Learning from Verifiable Rewards (RLVR) has substantially improved reasoning performance. However, exi…
cs.AI2024
Large Language Model Safety: A Holistic Survey
Dan Shi, Tianhao Shen, Yufei Huang +10
The rapid development and deployment of large language models (LLMs) have introduced a new frontier in artificial intelligence, marked by unprecedented capabilities in natural lang…