1 paper · 1 filter
Ziliang Wang, Kang An, Faqiang Qian +5
Although reinforcement learning (RL) has expanded the cognitive boundaries of large language models (LLMs), it often remains vulnerable to the autoregressive curse in long-horizon…