7 citations · 7 across the 4 of their papers we have counts for
1 paper · 1 filter
Skylar Zhai, Jingcheng Liang, Dongyeop Kang
Reinforcement fine-tuning improves the reasoning ability of large language models, but it can also encourage them to answer unanswerable queries by guessing or hallucinating missin…