1 paper · 1 filter
Shresth Verma, Mauricio Tec, Cheol Woo Kim +2
While LLMs excel at single-turn generation, they struggle with long-horizon, multi-turn interactions. Offline reinforcement learning (RL) offers a scalable approach, yet its perfor…