Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical Reasoning
Haolong Qian, Xianliang Yang, Yinuo ma +6
Knowledge distillation from powerful reasoning models is widely used to improve Small Language Models (SLMs) on mathematical reasoning, often assuming that traces with higher rewar…
cs.AI2026
DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations
Lirong Che, Yuzhe yang, Peiwen lin +3
Agent harness evolution improves frozen language-model agents by modifying the executable structures around them. We study this paradigm as a form of sample-efficient fast adaptati…