1 paper
Minhan Cho, Jimin Kweon
We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across domains and models (RPC across four new tas…