6 citations · 8 across the 6 of their papers we have counts for
7 papers
Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning
Tianyuan Shi, Canbin Huang, Bei Li +4
Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectively transferring what to answer rather th…
When Model Merging Breaks Routing: Training-Free Calibration for MoE
Canbin Huang, Tianyuan Shi, Xiaojun Quan +3
Model merging has emerged as a cost-effective approach for consolidating the capabilities of multiple LLMs without retraining. However, existing merging techniques, largely based o…
Lookahead Routing for Large Language Models
Canbin Huang, Tianyuan Shi, Yuhua Zhu +2
Large language model (LLM) routers improve the efficiency of multi-model systems by directing each query to the most appropriate model while leveraging the diverse strengths of het…
Mutual-Taught for Co-adapting Policy and Reward Models
Tianyuan Shi, Canbin Huang, Fanqi Wan +5
During the preference optimization of large language models (LLMs), distribution shifts may arise between newly generated model samples and the data used to train the reward model…
FuseRL: Dense Preference Optimization for Heterogeneous Model Fusion
Longguang Zhong, Fanqi Wan, Ziyi Yang +3
Heterogeneous model fusion enhances the performance of LLMs by integrating the knowledge and capabilities of multiple structurally diverse models. However, existing approaches ofte…
Weighted-Reward Preference Optimization for Implicit Model Fusion
Ziyi Yang, Fanqi Wan, Longguang Zhong +2
While fusing heterogeneous open-source LLMs with varying architectures and sizes can potentially integrate the strengths of different models, existing fusion methods face significa…