activity
20242026
collaborators

5 papers

cs.AI2026

The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical Reasoning

Haolong Qian, Xianliang Yang, Yinuo ma +6

Knowledge distillation from powerful reasoning models is widely used to improve Small Language Models (SLMs) on mathematical reasoning, often assuming that traces with higher rewar…

cs.LG2026

Holdout-Loss-Based Data Selection for LLM Finetuning via In-Context Learning

Ling Zhang, Xianliang Yang, Juwon Yu +4

Fine-tuning large pretrained language models is a common approach for aligning them with human preferences, but noisy or off-target examples can dilute supervision. While small, we…

cs.AI2025

HeurAgenix: Leveraging LLMs for Solving Complex Combinatorial Optimization Challenges

Xianliang Yang, Ling Zhang, Haolong Qian +2

Heuristic algorithms play a vital role in solving combinatorial optimization (CO) problems, yet traditional designs depend heavily on manual expertise and struggle to generalize ac…

cs.AI2024

Position: Rethinking Post-Hoc Search-Based Neural Approaches for Solving Large-Scale Traveling Salesman Problems

Yifan Xia, Xianliang Yang, Zichuan Liu +3

Recent advancements in solving large-scale traveling salesman problems (TSP) utilize the heatmap-guided Monte Carlo tree search (MCTS) paradigm, where machine learning (ML) models…

cs.MA2024

Knowing What Not to Do: Leverage Language Model Insights for Action Space Pruning in Multi-agent Reinforcement Learning

Zhihao Liu, Xianliang Yang, Zichuan Liu +7

Multi-agent reinforcement learning (MARL) is employed to develop autonomous agents that can learn to adopt cooperative or competitive strategies within complex environments. Howeve…