2 papers
cs.LG2026
AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage Margin
Jian Xiong, Jingbo Zhou, Jingyong Ye +2
Reinforcement learning (RL) has emerged as an effective approach for enhancing the reasoning capabilities of large language models (LLMs), especially in scenarios where supervised…
cs.LG2026
Unrewarded Exploration in Large Language Models Reveals Latent Learning from Psychology
Jian Xiong, Jingbo Zhou, Zihan Zhou +6
Latent learning, classically theorized by Tolman, shows that biological agents (e.g., rats) can acquire internal representations of their environment without rewards, enabling rapi…