1 paper · 1 filter
Rujiao Long, Yang Li, Xingyao Zhang +7
Exploration capacity shapes both inference-time performance and reinforcement learning (RL) training for large (vision-) language models, as stochastic sampling often yields redund…