4 papers
The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical Reasoning
Haolong Qian, Xianliang Yang, Yinuo ma +6
Knowledge distillation from powerful reasoning models is widely used to improve Small Language Models (SLMs) on mathematical reasoning, often assuming that traces with higher rewar…
RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents
Haitian Zhong, Jixiu Zhai, Lei Song +3
Multi-turn tool calling is challenging for Large Language Models (LLMs) because rewards are sparse and exploration is expensive. A common recipe, SFT followed by GRPO, can stall wh…
Holdout-Loss-Based Data Selection for LLM Finetuning via In-Context Learning
Ling Zhang, Xianliang Yang, Juwon Yu +4
Fine-tuning large pretrained language models is a common approach for aligning them with human preferences, but noisy or off-target examples can dilute supervision. While small, we…
HeurAgenix: Leveraging LLMs for Solving Complex Combinatorial Optimization Challenges
Xianliang Yang, Ling Zhang, Haolong Qian +2
Heuristic algorithms play a vital role in solving combinatorial optimization (CO) problems, yet traditional designs depend heavily on manual expertise and struggle to generalize ac…