Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
Zhi Zheng, Rongsheng Chen, Yunpeng Ba +3
Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rew…
cs.LG2026
Rethinking the Divergence Regularization in LLM RL
Jiarui Yao, Xiangxin Zhou, Penghui Qi +3
Reinforcement learning (RL) has become a key component of post-training large language models (LLMs). In practice, LLM RL is often off-policy because of training-inference mismatch…
cs.LG2025
Extending Epistemic Uncertainty Beyond Parameters Would Assist in Designing Reliable LLMs
T. Duy Nguyen-Hien, Desi R. Ivanova, Yee Whye Teh +1
Although large language models (LLMs) are highly interactive and extendable, current approaches to ensure reliability in deployments remain mostly limited to rejecting outputs with…