5 papers
Diverse Thinking Schemata Elicit Better Reasoning in Large Language Models
Xinyue Liang, Yizhe Yang, Yu Bai +3
Large reasoning models (LRMs) have attracted increasing attention for their ability to solve complex mathematical problems by generating extended reasoning chains. In this work, we…
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
Gao Yang, Yuhang Liu, Siyu Miao +3
Ideal or real - that is the question.In this work, we explore whether principles from game theory can be effectively applied to the evaluation of large language models (LLMs). This…
EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
Bin Xu, Yu Bai, Huashan Sun +10
As large language models continue to advance, their application in educational contexts remains underexplored and under-optimized. In this paper, we address this gap by introducing…
SEA: Semantic Map Prediction for Active Exploration of Uncertain Areas
Hongyu Ding, Xinyue Liang, Yudong Fang +7
In this paper, we propose SEA, a novel approach for active robot exploration through semantic map prediction and a reinforcement learning-based hierarchical exploration policy. Unl…
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment
Jiawei Li, Xinyue Liang, Junlong Zhang +3
Process supervision enhances the performance of large language models in reasoning tasks by providing feedback at each step of chain-of-thought reasoning. However, due to the lack…