3 citations · 3 across the 1 of their papers we have counts for
2 papers
cs.CL2025
SSPO: Subsentence-level Policy Optimization
Kun Yang, Zikang chen, Yanmeng Wang +4
As a key component of large language model (LLM) post-training, Reinforcement Learning from Verifiable Rewards (RLVR) has substantially improved reasoning performance. However, exi…
cs.AI2024★ 3 cited
Large Language Model Safety: A Holistic Survey
Dan Shi, Tianhao Shen, Yufei Huang +10
The rapid development and deployment of large language models (LLMs) have introduced a new frontier in artificial intelligence, marked by unprecedented capabilities in natural lang…