2 papers
cs.CL2026
BLADE: Boundary-Expanded and Layer-Adaptive Dynamic Exit for Efficient LLM Reasoning
Keshu Fu, Keqin Peng, Jun Bai +6
Large language models often improve task performance by generating long reasoning traces, but the resulting computation is frequently wasted on redundant verification and revision.…
cs.CL2026
Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation
Keqin Peng, Chen Li, Yuanxin Ouyang +2
On-Policy Self-Distillation (OPSD) has emerged as a crucial paradigm for enhancing and aligning Large Language Models (LLMs). However, in complex reasoning tasks, OPSD paradoxicall…