1 paper
Zhengbo Jiao, Hongyu Xian, Qinglong Wang +5
Large language models (LLMs) struggle with complex, long-horizon reasoning due to instability caused by their frozen policy assumption. Current test-time scaling methods treat exec…