1 citations · 1 across the 6 of their papers we have counts for
1 paper · 1 filter
Zihao Feng, Xiaoxue Wang, Bowen Wu +4
While reinforcement learning (RL) is increasingly used for LLM-based tool learning, its efficiency is often hampered by an overabundance of simple samples that provide diminishing…