2 papers
cs.CL2025
GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Jian Zhao, Runze Liu, Kaiyan Zhang +8
Recent advancements in Large Language Models (LLMs) have shown that it is promising to utilize Process Reward Models (PRMs) as verifiers to enhance the performance of LLMs. However…
cs.NE2024
Evolution of Thought: Diverse and High-Quality Reasoning via Multi-Objective Optimization
Biqing Qi, Zhouyi Qian, Yiang Luo +4
As multi-modal large language models (MLLMs) are increasingly applied to complex reasoning tasks, the diversity and quality of reasoning paths become crucial factors affecting thei…