4 papers
AdverMCTS: Combating Pseudo-Correctness in Code Generation via Adversarial Monte Carlo Tree Search
Qingyao Li, Weiwen Liu, Weinan Zhang +2
Recent advancements in Large Language Models (LLMs) have successfully employed search-based strategies to enhance code generation. However, existing methods typically rely on stati…
Conditional Performance Guarantee for Large Reasoning Models
Jianguo Huang, Hao Zeng, Bingyi Jing +2
Large reasoning models have shown strong performance through extended chain-of-thought reasoning, yet their computational cost remains significant. Probably approximately correct (…
Harnessing Reasoning Trajectories for Hallucination Detection via Answer-agreement Representation Shaping
Jianxiong Zhang, Bing Guo, Yuming Jiang +3
Large reasoning models (LRMs) often generate long, seemingly coherent reasoning traces yet still produce incorrect answers, making hallucination detection challenging. Although tra…
On the Provable Performance Guarantee of Efficient Reasoning Models
Hao Zeng, Jianguo Huang, Bingyi Jing +2
Large reasoning models (LRMs) have achieved remarkable progress in complex problem-solving tasks. Despite this success, LRMs typically suffer from high computational costs during d…