2 papers
cs.CL2026
Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models
Dan Shi, Zhuowen Han, Simon Ostermann +3
Reinforcement learning (RL)-based post-training often improves the reasoning performance of large language models (LLMs) beyond the training domain, while supervised fine-tuning (S…
cs.AI2024
Large Language Model Safety: A Holistic Survey
Dan Shi, Tianhao Shen, Yufei Huang +10
The rapid development and deployment of large language models (LLMs) have introduced a new frontier in artificial intelligence, marked by unprecedented capabilities in natural lang…