exploration-exploitation 1large language models 1overfitting mitigation 1skill learning 1transfer learning 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.AI2026
Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting
Hongqiang Lin, Chao Liu, Xiaofan Bai +4
The paper introduces SkillBoost, a three-stage framework that reduces overfitting of trainable skills in large language model agents by balancing constrained exploitation of failur…
cs.AI2026
Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief
Hongqiang Lin, Pengfei Wang, Nenggan Zheng
Offline reinforcement learning (RL) aims to optimize policies from pre-collected datasets. A bottleneck of this paradigm is managing epistemic uncertainty, which arises from limite…