2 papers
cs.AI2026
ICRL: Learning to Internalize Self-Critique with Reinforcement Learning
Jianbo Lin, Xiaomin Yu, Yi Xin +7
Large language model-based agents make mistakes, yet critique can often guide the same model toward correct behavior. However, when critique is removed, the model may fail again on…
cs.CL2023
CAME: Confidence-guided Adaptive Memory Efficient Optimization
Yang Luo, Xiaozhe Ren, Zangwei Zheng +3
Adaptive gradient methods, such as Adam and LAMB, have demonstrated excellent performance in the training of large language models. Nevertheless, the need for adaptivity requires m…