2 citations · 2 across the 3 of their papers we have counts for
1 paper · 2 filters
Wanli Yang, Hongyu Zang, Junwei Zhang +5
Reinforcement learning (RL) has achieved remarkable success in LLM reasoning, but whether it can also improve direct recall of parametric knowledge remains an open question. We stu…