2 papers
cs.AI2025
Hybrid Reward Normalization for Process-supervised Non-verifiable Agentic Tasks
Peiran Xu, Zhuohao Li, Xiaoying Xing +3
Large Language Models (LLMs) increasingly rely on external tools such as search engines to solve complex agentic tasks that require reasoning and external knowledge retrieval. Rece…
cs.AI2025
Hawkeye:Efficient Reasoning with Model Collaboration
Jianshu She, Zhuohao Li, Zhemin Huang +4
Chain-of-Thought (CoT) reasoning has demonstrated remarkable effectiveness in enhancing the reasoning abilities of large language models (LLMs). However, its efficiency remains a c…