Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
ChainPrune: Evaluating and Reducing Redundancy in Long Chain-of-Thought Reasoning
Weihang Pan, Zhengxu Yu, Yuxiang Zhang +5
Chain-of-Thought (CoT) reasoning has significantly enhanced the multi-step problem-solving capabilities of large language models (LLMs) by introducing explicit intermediate reasoni…
cs.LG2026
SPAR: Support-Preserving Action Rectification
Jiaxin Zhao, Weihang Pan, Xun Liang +1
Offline policy improvement faces an inherent conflict between maximizing value and fitting the data distribution. While in-sample weighted regression is stable, it suffers from ove…