3 papers
cs.AI2026
GR-Ben: A General Reasoning Benchmark for Evaluating Process Reward Models
Zhouhao Sun, Xuan Zhang, Xiao Ding +10
Currently, process reward models (PRMs) have exhibited remarkable potential for test-time scaling. Since large language models (LLMs) regularly generate flawed intermediate reasoni…
cs.CL2026
Large Language Models Are Still Misled by Simple Bias Ensembles
Zhouhao Sun, Zhiyuan Kan, Xiao Ding +5
With the evolution of large language models (LLMs), their robustness against individual simple biases has been enhanced. However, we observe that the ensemble of multiple simple bi…
cs.AI2025
Advances in Large Language Models for Medicine
Zhiyu Kan, Wensheng Gan, Zhenlian Qi +1
Artificial intelligence (AI) technology has advanced rapidly in recent years, with large language models (LLMs) emerging as a significant breakthrough. LLMs are increasingly making…