Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards
Li Wang, Xiaodong Lu, Xiaohan Wang +5
Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewards (RLVR). Nonetheless, RLVR i…
cs.LG2025
An Investigation of Batch Normalization in Off-Policy Actor-Critic Algorithms
Li Wang, Sudun, Xingjian Zhang +2
Batch Normalization (BN) has played a pivotal role in the success of deep learning by improving training stability, mitigating overfitting, and enabling more effective optimization…