5 papers
Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning
Tianyuan Shi, Canbin Huang, Bei Li +4
Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectively transferring what to answer rather th…
When Model Merging Breaks Routing: Training-Free Calibration for MoE
Canbin Huang, Tianyuan Shi, Xiaojun Quan +3
Model merging has emerged as a cost-effective approach for consolidating the capabilities of multiple LLMs without retraining. However, existing merging techniques, largely based o…
Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization
Shiping Gao, Hongzhan Chen, Xiaojun Quan +2
Process reward models (PRMs) provide fine-grained supervision for reasoning, but reliable PRMs often require step annotations or heavy verification pipelines, making them costly to…
Probabilistic Token Alignment for Large Language Model Fusion
Runjia Zeng, James Chenhao Liang, Cheng Han +8
Training large language models (LLMs) from scratch can yield models with unique functionalities and strengths, but it is costly and often leads to redundant capabilities. A more co…
PsyPlay: Personality-Infused Role-Playing Conversational Agents
Tao Yang, Yuhua Zhu, Xiaojun Quan +2
The current research on Role-Playing Conversational Agents (RPCAs) with Large Language Models (LLMs) primarily focuses on imitating specific speaking styles and utilizing character…