6 citations · 8 across the 8 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
When Model Merging Breaks Routing: Training-Free Calibration for MoE
Canbin Huang, Tianyuan Shi, Xiaojun Quan +3
Model merging has emerged as a cost-effective approach for consolidating the capabilities of multiple LLMs without retraining. However, existing merging techniques, largely based o…
cs.LG2025
Mutual-Taught for Co-adapting Policy and Reward Models
Tianyuan Shi, Canbin Huang, Fanqi Wan +5
During the preference optimization of large language models (LLMs), distribution shifts may arise between newly generated model samples and the data used to train the reward model…