3 papers
cs.CL2025
Entropy-Guided Reasoning Compression
Hourun Zhu, Yang Gao, Wenlong Fei +2
Large reasoning models have demonstrated remarkable performance on complex reasoning tasks, yet the excessive length of their chain-of-thought outputs remains a major practical bot…
cs.AI2025
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment
Jiawei Li, Xinyue Liang, Junlong Zhang +3
Process supervision enhances the performance of large language models in reasoning tasks by providing feedback at each step of chain-of-thought reasoning. However, due to the lack…
cs.LG2024
METEOR: Evolutionary Journey of Large Language Models from Guidance to Self-Growth
Jiawei Li, Xiaoang Xu, Yang Gao
Model evolution enables learning from feedback to refine experiences and update skills, transforming models from having no domain knowledge to becoming domain experts. However, the…