1 paper · 1 filter
Zetian Sun, Dongfang Li, Baotian Hu +2
In the Large Language Model(LLM) reasoning scenario, people often estimate state value via Monte Carlo sampling. Though Monte Carlo estimation is an elegant method with less induct…