6 papers
Diverse Thinking Schemata Elicit Better Reasoning in Large Language Models
Xinyue Liang, Yizhe Yang, Yu Bai +3
Large reasoning models (LRMs) have attracted increasing attention for their ability to solve complex mathematical problems by generating extended reasoning chains. In this work, we…
Entropy-Guided Reasoning Compression
Hourun Zhu, Yang Gao, Wenlong Fei +2
Large reasoning models have demonstrated remarkable performance on complex reasoning tasks, yet the excessive length of their chain-of-thought outputs remains a major practical bot…
Unveiling and Addressing Pseudo Forgetting in Large Language Models
Huashan Sun, Yizhe Yang, Yinghao Li +2
Although substantial efforts have been made to mitigate catastrophic forgetting in continual learning, the intrinsic mechanisms are not well understood. In this work, we demonstrat…
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment
Jiawei Li, Xinyue Liang, Junlong Zhang +3
Process supervision enhances the performance of large language models in reasoning tasks by providing feedback at each step of chain-of-thought reasoning. However, due to the lack…
METEOR: Evolutionary Journey of Large Language Models from Guidance to Self-Growth
Jiawei Li, Xiaoang Xu, Yang Gao
Model evolution enables learning from feedback to refine experiences and update skills, transforming models from having no domain knowledge to becoming domain experts. However, the…
PSST: A Benchmark for Evaluation-driven Text Public-Speaking Style Transfer
Huashan Sun, Yixiao Wu, Yuhao Ye +4
Language style is necessary for AI systems to understand and generate diverse human language accurately. However, previous text style transfer primarily focused on sentence-level d…