1 paper · 1 filter
Guangran Cheng, Chengqi Lyu, Songyang Gao +2
Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), helping balance exploration-exploitation trade-off during alignment pro…