Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
SAPO: Self-Adaptive Process Optimization Makes Small Reasoners Stronger
Kaiyuan Chen, Guangmin Zheng, Jin Wang +2
Existing self-evolution methods overlook the influence of fine-grained reasoning steps, which leads to the reasoner-verifier gap. The computational inefficiency of Monte Carlo (MC)…
cs.CL2024
Learning to Reason via Self-Iterative Process Feedback for Small Language Models
Kaiyuan Chen, Jin Wang, Xuejie Zhang
Small language models (SLMs) are more efficient, cost-effective, and customizable than large language models (LLMs), though they often underperform in specific areas like reasoning…