1 paper
Hua-Dong Xiong, Xinyuan Yan, Ji-An Li +3
Inference-time thinking improves the performance of large language models, but aggregate outcomes do not reveal whether models use available evidence more effectively or seek infor…