1 paper · 1 filter
Shengxuan Qiu, Haochen Huang, Shuzhang Zhong +2
Scaling test-time compute with multi-path chain-of-thought improves reasoning accuracy, but its effectiveness depends critically on the exploration-exploitation trade-off. Existing…