Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making
Guangfeng Cai, Kaibing Yang, Shuo He +4
Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties super…
cs.CL2026
Online Causal Kalman Filtering for Stable and Effective Policy Optimization
Shuo He, Lang Feng, Xin Cheng +2
Reinforcement learning for large language models suffers from high-variance token-level importance sampling (IS) ratios, which would destabilize policy optimization at scale. To im…