3 papers
cs.DC2025
EC2MoE: Adaptive End-Cloud Pipeline Collaboration Enabling Scalable Mixture-of-Experts Inference
Zheming Yang, Yunqing Hu, Sheng Sun +1
The Mixture-of-Experts (MoE) paradigm has emerged as a promising solution to scale up model capacity while maintaining inference efficiency. However, deploying MoE models across he…
cs.LG2025
Adversarial Training of Reward Models
Alexander Bukharin, Haifeng Qian, Shengyang Sun +6
Reward modeling has emerged as a promising approach for the scalable alignment of language models. However, contemporary reward models (RMs) often lack robustness, awarding high re…
cs.LG2025
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment
Shengyang Sun, Yian Zhang, Alexander Bukharin +11
The rapid development of large language model (LLM) alignment algorithms has resulted in a complex and fragmented landscape, with limited clarity on the effectiveness of different…