5 papers
Bayesian Preference Learning for Test-Time Steerable Reward Models
Jiwoo Hong, Shao Tang, Zhipeng Wang
Reward models are central to aligning language models with human preferences via reinforcement learning (RL). As RL is increasingly applied to settings such as verifiable rewards a…
Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
Xiaofeng Lin, Sirou Zhu, Yilei Chen +6
Large language models (LLMs) achieve strong performance when all task-relevant information is available upfront, as in static prediction and instruction-following problems. However…
Aligning Diffusion Language Models via Unpaired Preference Optimization
Vaibhav Jindal, Hejian Sang, Chun-Mao Lai +2
Diffusion language models (dLLMs) are an emerging alternative to autoregressive (AR) generators, but aligning them to human preferences is challenging because sequence log-likeliho…
Scaling Up Efficient Small Language Models Serving and Deployment for Semantic Job Search
Kayhan Behdin, Qingquan Song, Sriram Vasudevan +17
Large Language Models (LLMs) have demonstrated impressive quality when applied to predictive tasks such as relevance ranking and semantic search. However, deployment of such LLMs r…
Debunk the Myth of SFT Generalization
Xiaofeng Lin, Hejian Sang, Zhipeng Wang +1
A prevailing view holds that supervised fine-tuning (SFT) memorizes training data and fails to generalize, whereas reinforcement learning (RL) attains broader robustness. We revisi…