1 citations · 1 across the 2 of their papers we have counts for
4 papers
UniARM: Towards a Unified Autoregressive Reward Model for Multi-Objective Test-Time Alignment
Hongyan Xie, Yikun Ban, Ruiyu Fang +6
Multi-objective alignment aims to align LLM responses with multiple human preference objectives. Among existing methods, guiding the generation of frozen LLMs through autoregressiv…
Your Group-Relative Advantage Is Biased
Fengkai Yang, Zherui Chen, Xiaohan Wang +10
Reinforcement Learning from Verifier Rewards (RLVR) has emerged as a widely used approach for post-training large language models on reasoning tasks, with group-based methods such…
LLMBoost: Make Large Language Models Stronger with Boosting
Zehao Chen, Tianxiang Ai, Yifei Li +11
Ensemble learning of LLMs has emerged as a promising alternative to enhance performance, but existing approaches typically treat models as black boxes, combining the inputs or fina…
Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
Zixuan Huang, Yikun Ban, Lean Fu +4
Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its performance is highly depen…