3 papers
cs.LG2025
On Predictability of Reinforcement Learning Dynamics for Large Language Models
Yuchen Cai, Ding Cao, Xin Xu +7
Recent advances in reasoning capabilities of large language models (LLMs) are largely driven by reinforcement learning (RL), yet the underlying parameter dynamics during RL trainin…
cs.CL2025
On the Superimposed Noise Accumulation Problem in Sequential Knowledge Editing of Large Language Models
Ding Cao, Yuchen Cai, Yuqing Huang +4
Sequential knowledge editing techniques aim to continuously update knowledge in large language models at low cost, preventing models from generating outdated or incorrect informati…
cs.CL2024
O-Edit: Orthogonal Subspace Editing for Language Model Sequential Editing
Yuchen Cai, Ding Cao
Large language models (LLMs) acquire knowledge during pre-training, but over time, this knowledge may become incorrect or outdated, necessitating updates after training. Knowledge…