4 citations · 4 across the 8 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces
Chen He, Yuhao Wu, Lei Wang +2
Long chain-of-thought (CoT) traces are widely used as supervision for reasoning-oriented LLM SFT, yet answer-correct traces can still lead to markedly different fine-tuning outcome…
cs.AI2025
What Makes Reasoning Invalid: Echo Reflection Mitigation for Large Language Models
Chen He, Xun Jiang, Lei Wang +5
Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of reasoning tasks. Recent methods have further improved LLM performance in complex mathem…