2 citations · 2 across the 11 of their papers we have counts for
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces
Chen He, Yuhao Wu, Lei Wang +2
Long chain-of-thought (CoT) traces are widely used as supervision for reasoning-oriented LLM SFT, yet answer-correct traces can still lead to markedly different fine-tuning outcome…
cs.AI2026★ 2 cited
SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit Assignment
Hongxiang Lin, Zhirui Kuai, Erpeng Xue +1
Structured skill prompts improve exploration in long-horizon agentic reinforcement learning (RL). Skill-augmented RL methods retain external skills at inference, while skill-intern…
cs.AI2026
Scientific Logicality Enriched Methodology for LLM Reasoning: A Practice in Physics
Zhaoxin Yu, Nan Xu, Kun Chen +3
With the continuous advancement of reasoning abilities in Large Language Models (LLMs), their application to scientific reasoning tasks has gained significant research attention. C…